Abstract
Knowledge distillation (KD) has become a key technique for compressing object detection models while maintaining accuracy, enabling deployment in latency-sensitive scenarios such as autonomous driving and drone surveillance. However, existing detection-oriented KD approaches face two major limitations: first, attention modules transplanted from classification tasks fail to capture fine-grained localization cues and perform poorly under occlusion; second, the naive combination of feature imitation and response mimicking often leads to performance degradation and slower training. To address these issues, we propose a Unified Attention Distillation (UAD) framework that integrates masked local attention to enhance textures and edges with a Row-Column Masked Global Attention (RCMGA) module for long-range dependency modeling and occlusion-aware suppression. Furthermore, we introduce a Decoupled Progressive Schedule (DPS) that progressively transitions from unified to global attention in terms of spatial granularity, and from feature imitation to response mimicking in terms of knowledge hierarchy, thereby harmonizing the dynamics of knowledge transfer. Extensive experiments on COCO and PASCAL VOC demonstrate that UAD combined with DPS achieves state-of-the-art detection performance while reducing training time by 11.9% compared to existing distillation methods.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- decoupled progressive schedule
- knowledge distillation
- Object detection
- unified attention distillation
Fingerprint
Dive into the research topics of 'Unified Attention Distillation with Decoupled Progressive Scheduling for Object Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver