TY - JOUR
T1 - Domain-adaptive Class Prototype Learning for Domain Generalization
AU - Wang, Shuyue
AU - Xu, Lian
AU - Boussaid, Farid
AU - Bennamoun, Mohammed
AU - Liu, Zhunga
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Domain Generalization (DG) aims to develop models that generalize effectively to unseen target domains despite significant distribution shifts. Pretrained vision-language models (VLMs), such as CLIP, have demonstrated strong generalization capabilities in DG tasks, largely due to their robust semantic representations. Recent approaches attempt to enhance CLIP by incorporating domain-specific information into textual class prompts, typically by leveraging visual domain features. However, such strategies may limit the ability to capture representative domain-specific semantics while maintaining strong generalization to unseen domains. In this work, we propose a novel framework that constructs domain-adaptive class prototypes by jointly optimizing domain and class prototypes within CLIP's vision-language aligned space, facilitating robust cross-domain classification. A core component of our framework is the text-aligned domain prototype learning module, which aligns visual domain prototypes with domain-aware textual embeddings to capture domain-specific semantics that are both representative and generalizable. These aligned domain prototypes are then integrated with class-level semantics to adapt class representations to each domain. Additionally, we incorporate a domain-invariant visual classifier to complement predictions from domain-adaptive class prototypes, enhancing stability and robustness under visual distribution shifts. Extensive experiments on five widely used DG benchmarks demonstrate the superiority of our method.
AB - Domain Generalization (DG) aims to develop models that generalize effectively to unseen target domains despite significant distribution shifts. Pretrained vision-language models (VLMs), such as CLIP, have demonstrated strong generalization capabilities in DG tasks, largely due to their robust semantic representations. Recent approaches attempt to enhance CLIP by incorporating domain-specific information into textual class prompts, typically by leveraging visual domain features. However, such strategies may limit the ability to capture representative domain-specific semantics while maintaining strong generalization to unseen domains. In this work, we propose a novel framework that constructs domain-adaptive class prototypes by jointly optimizing domain and class prototypes within CLIP's vision-language aligned space, facilitating robust cross-domain classification. A core component of our framework is the text-aligned domain prototype learning module, which aligns visual domain prototypes with domain-aware textual embeddings to capture domain-specific semantics that are both representative and generalizable. These aligned domain prototypes are then integrated with class-level semantics to adapt class representations to each domain. Additionally, we incorporate a domain-invariant visual classifier to complement predictions from domain-adaptive class prototypes, enhancing stability and robustness under visual distribution shifts. Extensive experiments on five widely used DG benchmarks demonstrate the superiority of our method.
KW - Domain generalization
KW - domain-specific semantics
KW - image-text alignment
KW - vision-language models (VLMs)
UR - https://www.scopus.com/pages/publications/105045327580
U2 - 10.1109/TMM.2026.3714585
DO - 10.1109/TMM.2026.3714585
M3 - 文章
AN - SCOPUS:105045327580
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -