Abstract
Domain Generalization (DG) aims to develop models that generalize effectively to unseen target domains despite significant distribution shifts. Pretrained vision-language models (VLMs), such as CLIP, have demonstrated strong generalization capabilities in DG tasks, largely due to their robust semantic representations. Recent approaches attempt to enhance CLIP by incorporating domain-specific information into textual class prompts, typically by leveraging visual domain features. However, such strategies may limit the ability to capture representative domain-specific semantics while maintaining strong generalization to unseen domains. In this work, we propose a novel framework that constructs domain-adaptive class prototypes by jointly optimizing domain and class prototypes within CLIP's vision-language aligned space, facilitating robust cross-domain classification. A core component of our framework is the text-aligned domain prototype learning module, which aligns visual domain prototypes with domain-aware textual embeddings to capture domain-specific semantics that are both representative and generalizable. These aligned domain prototypes are then integrated with class-level semantics to adapt class representations to each domain. Additionally, we incorporate a domain-invariant visual classifier to complement predictions from domain-adaptive class prototypes, enhancing stability and robustness under visual distribution shifts. Extensive experiments on five widely used DG benchmarks demonstrate the superiority of our method.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Domain generalization
- domain-specific semantics
- image-text alignment
- vision-language models (VLMs)
Fingerprint
Dive into the research topics of 'Domain-adaptive Class Prototype Learning for Domain Generalization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver