摘要
Knowledge distillation (KD) has emerged as a powerful technique for transferring knowledge from large, complex teacher models to smaller, more efficient student models. However, current KD methods primarily concentrate on mimicking instance-level predictions or feature representations, often overlooking the crucial role of class-level semantic structure in guiding effective knowledge transfer. This paper introduces Prototypical Decoupled Knowledge Distillation (PDKD), a novel framework designed to bridge the gap between instance-specific and class-discriminative knowledge by leveraging class prototypes. PDKD incorporates a prototype-aware supervision module that distills global class characteristics by aligning student predictions with both instance-level and prototype-based outputs from the teacher. This module dynamically harmonizes the logit scales of these two targets, effectively addressing the model size mismatches between teacher and student. Furthermore, a feature discrepancy alignment module is proposed to enforce consistency between the teacher and student in how they modulate features between the learned prototypes and individual samples. This alignment preserves the structural relationships between classes. By effectively unifying hierarchical class semantics with instance-level learning, PDKD establishes a new paradigm for training compact yet highly discriminative models. Extensive experiments on CIFAR-100 and ImageNet showcase the superior performance of PDKD compared to existing state-of-the-art methods.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 7979-7991 |
| 页数 | 13 |
| 期刊 | IEEE Transactions on Circuits and Systems for Video Technology |
| 卷 | 36 |
| 期 | 6 |
| DOI | |
| 出版状态 | 已出版 - 1 6月 2026 |
学术指纹
探究 'Prototype Decoupled Knowledge Distillation' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver