TY - JOUR
T1 - Unified Attention Distillation with Decoupled Progressive Scheduling for Object Detection
AU - Song, Yaoye
AU - Zhang, Peng
AU - Zhang, Yanning
AU - Zheng, Yefeng
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Knowledge distillation (KD) has become a key technique for compressing object detection models while maintaining accuracy, enabling deployment in latency-sensitive scenarios such as autonomous driving and drone surveillance. However, existing detection-oriented KD approaches face two major limitations: first, attention modules transplanted from classification tasks fail to capture fine-grained localization cues and perform poorly under occlusion; second, the naive combination of feature imitation and response mimicking often leads to performance degradation and slower training. To address these issues, we propose a Unified Attention Distillation (UAD) framework that integrates masked local attention to enhance textures and edges with a Row-Column Masked Global Attention (RCMGA) module for long-range dependency modeling and occlusion-aware suppression. Furthermore, we introduce a Decoupled Progressive Schedule (DPS) that progressively transitions from unified to global attention in terms of spatial granularity, and from feature imitation to response mimicking in terms of knowledge hierarchy, thereby harmonizing the dynamics of knowledge transfer. Extensive experiments on COCO and PASCAL VOC demonstrate that UAD combined with DPS achieves state-of-the-art detection performance while reducing training time by 11.9% compared to existing distillation methods.
AB - Knowledge distillation (KD) has become a key technique for compressing object detection models while maintaining accuracy, enabling deployment in latency-sensitive scenarios such as autonomous driving and drone surveillance. However, existing detection-oriented KD approaches face two major limitations: first, attention modules transplanted from classification tasks fail to capture fine-grained localization cues and perform poorly under occlusion; second, the naive combination of feature imitation and response mimicking often leads to performance degradation and slower training. To address these issues, we propose a Unified Attention Distillation (UAD) framework that integrates masked local attention to enhance textures and edges with a Row-Column Masked Global Attention (RCMGA) module for long-range dependency modeling and occlusion-aware suppression. Furthermore, we introduce a Decoupled Progressive Schedule (DPS) that progressively transitions from unified to global attention in terms of spatial granularity, and from feature imitation to response mimicking in terms of knowledge hierarchy, thereby harmonizing the dynamics of knowledge transfer. Extensive experiments on COCO and PASCAL VOC demonstrate that UAD combined with DPS achieves state-of-the-art detection performance while reducing training time by 11.9% compared to existing distillation methods.
KW - decoupled progressive schedule
KW - knowledge distillation
KW - Object detection
KW - unified attention distillation
UR - https://www.scopus.com/pages/publications/105046277988
U2 - 10.1109/TMM.2026.3718631
DO - 10.1109/TMM.2026.3718631
M3 - 文章
AN - SCOPUS:105046277988
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -