Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation

Lili WU , Renmin ZHANG , Bin ZHANG , Jincheng ZHANG , Xi CHEN , Yingjing QIAN

Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (9) : 260160

PDF (815KB)
Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (9) :260160 DOI: 10.1631/ENG.ITEE.2026.0160
Research Article
Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation
Author information +
History +
PDF (815KB)

Abstract

Multimodal recommendation aims to enrich preference modeling by leveraging visual and textual features. However, integrating high-dimensional pretrained features introduces substantial computational overhead. While knowledge distillation provides an effective compression strategy, existing frameworks face three intertwined challenges: rank bottlenecks caused by low-dimensional projections, cross-modal interference induced by shared fusion spaces, and optimization instability under static distillation temperatures. To address these issues, we propose ProMoE-DTS, a decoupled prompt-guided mixture-of-experts framework with dynamic temperature scheduling. Using an asymmetric teacher-student architecture, the teacher model leverages modality-aware soft prompts as semantic anchors to route heterogeneous features into parameter-disjoint expert networks, thereby alleviating cross-modal conflicts and resolving the rank bottlenecks. To ensure stable knowledge transfer, a feedback-driven dynamic temperature scheduler adaptively regulates the distillation intensity based on epoch-wise signals. This asymmetric design confines intensive multimodal operations to the offline teacher, leaving the online student model with a highly efficient, pure identifier-based structure. Extensive experiments on three benchmark datasets demonstrate that ProMoE-DTS improves Recall@20 by 2.24%-3.96% over state-of-the-art baselines, while requiring only 3.28%-3.55% of the teacher's parameters.

Keywords

Multimodal recommendation / Knowledge distillation (KD) / Prompt tuning / Mixture-of-experts (MoE) / Dynamic temperature scheduling

Cite this article

Download citation ▾
Lili WU, Renmin ZHANG, Bin ZHANG, Jincheng ZHANG, Xi CHEN, Yingjing QIAN. Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation. Eng Inform Technol Electron Eng, 2026, 27 (9) : 260160 DOI:10.1631/ENG.ITEE.2026.0160

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Bao KQ , Zhang JZ , Zhang Y , et al., 2023. TALLRec:an effective and efficient tuning framework to align large language model with recommendation. Proc 17th ACM Conf on Recommender Systems, p.1007- 1014.

[2]

Chen JY , Zhang HW , He XN , et al., 2017. Attentive collaborative filtering:multimedia recommendation with item- and componentlevel attention. Proc 40th Int ACM SIGIR Conf on Research and Development in Information Retrieval, p.335- 344.

[3]

Dai BQ , Du ZC , Zhu JM , et al., 2024. UniEmbedding:learning universal multi-modal multi-domain item embeddings via userview contrastive learning. Proc 33rd ACM Int Conf on Information and Knowledge Management, p.4446- 4453.

[4]

Deng JX , Wang SY , Cai K , et al., 2025. OneRec:unifying retrieve and rank with generative recommender and iterative preference alignment. https://arxiv.org/abs/2502.18965

[5]

Dong XM , Huang O , Thulasiraman P , et al., 2023. Improved knowledge distillation via teacher assistants for sentiment analysis. Proc IEEE Symp Series on Computational Intelligence, p.300- 305.

[6]

Du YP , Sun Z , Wang ZY , et al., 2025. Active large language modelbased knowledge distillation for session-based recommendation. Proc 39th AAAI Conf on Artificial Intelligence, p.11607- 11615.

[7]

Geng X , Zhang HW , Bian JW , et al., 2015. Learning image and user features for recommendation in social networks. Proc IEEE Int Conf on Computer Vision, p.4274- 4282.

[8]

He RN , McAuley J , 2016. VBPR:visual Bayesian personalized ranking from implicit feedback. Proc 30th AAAI Conf on Artificial Intelligence, p.144- 150.

[9]

He XN , Deng K , Wang X , et al., 2020. LightGCN:simplifying and powering graph convolution network for recommendation. Proc 43rd Int ACM SIGIR Conf on Research and Development in Information Retrieval, p.639- 648.

[10]

Hinton G , Vinyals O , Dean J , 2015. Distilling the knowledge in a neural network. https://arxiv.org/abs/1503.02531

[11]

Hu XH , Zhang HT , 2025. Invariant representation learning in multimedia recommendation with modality alignment and model fusion. Entropy, 27 (1): 56.

[12]

Huo FS , Xu WC , Guo JC , et al., 2024. C2KD:bridging the modality gap for cross-modal knowledge distillation. Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.16006- 16015.

[13]

Islam SU , Ahad JI , Rahman F , et al., 2025. Dynamic temperature scheduler for knowledge distillation. https://arxiv.org/abs/2511.13767

[14]

Jian M , Wang T , Yang MJ , et al., 2026. Hierarchy-aware multimodal distillation for recommendation. IEEE Trans Multim, 28: 2279- 2290.

[15]

Kang S , Kweon W , Lee D , et al., 2023. Distillation from heterogeneous models for top-K recommendation. Proc ACM Web Conf, p.801- 811.

[16]

Lan PX , Xu HY , Yang EN , et al., 2025. Efficient and effective prompt tuning via prompt decomposition and compressed outer product. Proc Conf of the Nations of the Americas Chapter of the Association for Computational Linguistics:Human Language Technologies, p.4406- 4421.

[17]

Lee JW , Choi M , Lee J , et al., 2019. Collaborative distillation for top-N recommendation. Proc IEEE Int Conf on Data Mining, p.369- 378.

[18]

Li L , Zhang YF , Chen L , 2023. Personalized prompt learning for explainable recommendation. ACM Trans Inform Syst, 41 (4): 103.

[19]

Liang JH , Zhao XY , Li MY , et al., 2023. MMMLP:multi-modal multilayer perceptron for sequential recommendations. Proc ACM Web Conf, p.1109- 1117.

[20]

Lin JH , Chen B , Wang HY , et al., 2024. ClickPrompt:CTR models are strong prompt generators for adapting language models to CTR prediction. Proc ACM Web Conf, p.3319- 3330.

[21]

Lin YF , He JY , 2026. A reinforcement learning framework for multimodal recommendation explanation generation. Authorea.

[22]

Liu F , Chen HL , Cheng ZY , et al., 2023. Disentangled multimodal representation learning for recommendation. IEEE Trans Multim, 25: 7149- 7159.

[23]

Liu MR , Zhang SX , Long C , 2025. Facet-aware multi-head mixtureof-experts model for sequential recommendation. Proc 18th ACM Int Conf on Web Search and Data Mining, p.127- 135.

[24]

Liu YF , Zhang KN , Ren XY , et al., 2024. AlignRec:aligning and training in multimodal recommendations. Proc 33rd ACM Int Conf on Information and Knowledge Management, p.1503- 1512.

[25]

Loshchilov I , Hutter F , 2019. Decoupled weight decay regularization.

[26]

Ma WY , Xia HB , Liu Y , 2025. DiffKD:collaborative graph diffusion with knowledge distillation for multimodal recommendation. J Intell Inform Syst, 63 (5): 1487- 1510.

[27]

McAuley J , Targett C , Shi QF , et al., 2015. Image-based recommendations on styles and substitutes. Proc 38th Int ACM SIGIR Conf on Research and Development in Information Retrieval, p.43- 52.

[28]

Mo F , Xiao L , Song QY , et al., 2025. FGCM:modality-behavior fusion model integrated with graph contrastive learning for multimodal recommendation. IEEE Multim, 32 (3): 27- 37.

[29]

Nguyen NH , Nguyen TA , Nguyen T , et al., 2024. Towards efficient communication and secure federated recommendation system via low-rank training. Proc ACM Web Conf, p.3940- 3951.

[30]

Qiu RH , Wang S , Chen Z , et al., 2021. CausalRec:causal inference for visual debiasing in visually-aware recommendation. Proc 29th ACM Int Conf on Multimedia, p.3844- 3852.

[31]

Reimers N , Gurevych I , 2019. Sentence-BERT:sentence embeddings using Siamese BERT-networks. Proc Conf on Empirical Methods in Natural Language Processing and the 9th Int Joint Conf on Natural Language Processing, p.3982- 3992.

[32]

Rendle S , Freudenthaler C , Gantner Z , et al., 2009. BPR:Bayesian personalized ranking from implicit feedback. Proc 25th Conf on Uncertainty in Artificial Intelligence, p.452- 461.

[33]

Tang JX , Wang K , 2018. Ranking distillation:learning compact ranking models with high performance for recommender system. Proc 24th ACM SIGKDD Int Conf on Knowledge Discovery and Data Mining, p.2289- 2298.

[34]

Tian YJ , Zhang CX , Guo ZC , et al., 2022. NOSMOG:learning noiserobust and structure-aware MLPs on graphs.

[35]

Wang J , Zhu L , Dai T , et al., 2021. Low-rank and sparse matrix factorization with prior relations for recommender systems. Appl Intell, 51 (6): 3435- 3449.

[36]

Wang QY , Yin HZ , Wang H , et al., 2019. Enhancing collaborative filtering with generative augmentation. Proc 25th ACM SIGKDD Int Conf on Knowledge Discovery and Data Mining, p.548- 556.

[37]

Wang X , He XN , Wang M , et al., 2019. Neural graph collaborative filtering. Proc 42nd Int ACM SIGIR Conf on Research and Development in Information Retrieval, p.165- 174.

[38]

Wei RX , Lan JM , Li KK , et al., 2025. MFDB:multimodal feature fusion and dynamic behavior modeling for interactive recommendation systems. Knowl-Based Syst, 326: 114047.

[39]

Wei W , Huang C , Xia LH , et al., 2023. Multi-modal self-supervised learning for recommendation. Proc ACM Web Conf, p.790- 800.

[40]

Wei W , Tang JB , Xia LH , et al., 2024. PromptMM:multi-modal knowledge distillation for recommendation with prompt-tuning. Proc ACM Web Conf, p.3217- 3228.

[41]

Wei YW , Wang X , Nie LQ , et al., 2019. MMGCN:multi-modal graph convolution network for personalized recommendation of microvideo. Proc 27th ACM Int Conf on Multimedia, p.1437- 1445.

[42]

Wei YW , Wang X , Nie LQ , et al., 2020. Graph-refined convolutional network for multimedia recommendation with implicit feedback. Proc 28th ACM Int Conf on Multimedia, p.3541- 3549.

[43]

Wei YW , Wang X , He XN , et al., 2022. Hierarchical user intent graph network for multimedia recommendation. IEEE Trans Multim, 24: 2701- 2712.

[44]

Xun JH , Zhang SY , Zhao Z , et al., 2021. Why do we click:visual impression-aware news recommendation. Proc 29th ACM Int Conf on Multimedia, p.3881- 3890.

[45]

Yu HF , Qi ZY , Jang LK , et al., 2024. MMoE:enhancing multimodal models with mixtures of multimodal interaction experts. Proc Conf on Empirical Methods in Natural Language Processing, p.10006- 10030.

[46]

Zhang JH , Zhu YQ , Liu Q , et al., 2021. Mining latent structures for multimedia recommendation. Proc 29th ACM Int Conf on Multimedia, p.3872- 3880.

[47]

Zhang ZJ , Liu SC , Yu JA , et al., 2024. M3oE:multi-domain multi-task mixture-of-experts recommendation framework. Proc 47th Int ACM SIGIR Conf on Research and Development in Information Retrieval, p.893- 902.

[48]

Zhou X , Zhou HY , Liu Y , et al., 2023. Bootstrap latent representations for multi-modal recommendation. Proc ACM Web Conf, p.845- 854.

[49]

Zhu NJ , Ren YQ , Liu Y , et al., 2025. MMGCL:multi-scale and multichannel graph contrastive learning for flight anomaly detection. Knowl-Based Syst, 329: 114275.

Rights & permissions

The Authors. Published by Zhejiang University Press Co., Ltd.

PDF (815KB)

Supplementary files

EITEE20260902-LLW-suppl2

6

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉