TY - JOUR
T1 - IPRec
T2 - A multimodal recommendation model with item-specific features and progressive knowledge distillation
AU - Feng, Junmei
AU - Zhao, Yaomin
AU - Zhang, Yihan
AU - Miao, Qiguang
AU - Lu, Zixiang
AU - Xia, Zhaoqiang
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2027/1/1
Y1 - 2027/1/1
N2 - Multimodal recommender systems have become essential to modern content platforms, leveraging user-item interactions to continually improve recommendation quality. However, many existing approaches still rely on coarse, globally shared fusion schemes and struggle to remain robust when multimodal content is incomplete. Consequently, we propose IPRec, a multimodal recommendation model, to tackle two persistent challenges: the lack of item-aware fusion strategies and the limited robustness of multimodal representations. IPRec mainly consists of two key components, i.e., Adaptive Item-specific Feature Learning (AIFL) and Progressive Knowledge Distillation Enhancement (PKDE). The AIFL module learns item-specific weight distributions across refined visual and textual representations from a large language model, yielding content-aware multimodal fusion weights that adapts to each item’s characteristics. The PKDE module enhances model robustness by combining masked auto-encoders with teacher-student distillation, allowing the student to effectively handle missing content while aligning with the teacher’s cross-modal representations. Extensive experiments on real-world datasets demonstrate that IPRec consistently outperforms the state-of-the-art methods, highlighting the importance of two components for enhancing multimodal recommenders.
AB - Multimodal recommender systems have become essential to modern content platforms, leveraging user-item interactions to continually improve recommendation quality. However, many existing approaches still rely on coarse, globally shared fusion schemes and struggle to remain robust when multimodal content is incomplete. Consequently, we propose IPRec, a multimodal recommendation model, to tackle two persistent challenges: the lack of item-aware fusion strategies and the limited robustness of multimodal representations. IPRec mainly consists of two key components, i.e., Adaptive Item-specific Feature Learning (AIFL) and Progressive Knowledge Distillation Enhancement (PKDE). The AIFL module learns item-specific weight distributions across refined visual and textual representations from a large language model, yielding content-aware multimodal fusion weights that adapts to each item’s characteristics. The PKDE module enhances model robustness by combining masked auto-encoders with teacher-student distillation, allowing the student to effectively handle missing content while aligning with the teacher’s cross-modal representations. Extensive experiments on real-world datasets demonstrate that IPRec consistently outperforms the state-of-the-art methods, highlighting the importance of two components for enhancing multimodal recommenders.
KW - Adaptive feature learning
KW - Knowledge distillation
KW - Multimodal recommendation
KW - Robust representation
UR - https://www.scopus.com/pages/publications/105045101346
U2 - 10.1016/j.eswa.2026.133632
DO - 10.1016/j.eswa.2026.133632
M3 - 文章
AN - SCOPUS:105045101346
SN - 0957-4174
VL - 332
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 133632
ER -