TY - JOUR
T1 - Adaptively Hard-Aware Temperature Scaling for Multi-Label Distillation
AU - Li, Cong
AU - Cheng, Gong
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - This article aims to fill the research gap of the biased learning problem in logit knowledge distillation (KD) under the multi-label learning (MLL) context. We construct relations between temperature scaling, a critical technique in logit KD, and two MLL inherent problems: positive-negative label imbalance and learning difficulty imbalance. With this perspective, we first uncover that the existential logit KD schemes, both theoretically, lead to insufficient distillation on the beneficial positive labels and, empirically, yield poor KD performance. In particular, we introduce the tempered sigmoid and demonstrate its hard-mining effect, with the tempered teacher-student predictive discrepancy governing the penalty strength per label. Building upon these findings, we further put forth a novel KD method dubbed AHTD, dynamically balancing the distilling effort put on hard-positive labels and otherwise. Incorporating both the ground-truth and the teacher’s and student’s predictions, AHTD is instantiated by: 1) an adaptive indicator to track the hard-positive labels along the distillation course; 2) a predictive discrepancy guided scaling strategy, which is aware of the hardness level regarding specific image, label, and distillation stage. Our work is accompanied by extensive experiments on MS-COCO, PASCAL-VOC, NUS-WIDE, OpenImage, and LVIS with multiple teacher-student pairs spanning across image classification, object detection, and instance segmentation. Distillation results demonstrate that AHTD consistently outperforms its previous counterparts by a clear margin and maintains a pleasing level of methodology simplicity and training efficiency, manifesting its superiority and scalability.
AB - This article aims to fill the research gap of the biased learning problem in logit knowledge distillation (KD) under the multi-label learning (MLL) context. We construct relations between temperature scaling, a critical technique in logit KD, and two MLL inherent problems: positive-negative label imbalance and learning difficulty imbalance. With this perspective, we first uncover that the existential logit KD schemes, both theoretically, lead to insufficient distillation on the beneficial positive labels and, empirically, yield poor KD performance. In particular, we introduce the tempered sigmoid and demonstrate its hard-mining effect, with the tempered teacher-student predictive discrepancy governing the penalty strength per label. Building upon these findings, we further put forth a novel KD method dubbed AHTD, dynamically balancing the distilling effort put on hard-positive labels and otherwise. Incorporating both the ground-truth and the teacher’s and student’s predictions, AHTD is instantiated by: 1) an adaptive indicator to track the hard-positive labels along the distillation course; 2) a predictive discrepancy guided scaling strategy, which is aware of the hardness level regarding specific image, label, and distillation stage. Our work is accompanied by extensive experiments on MS-COCO, PASCAL-VOC, NUS-WIDE, OpenImage, and LVIS with multiple teacher-student pairs spanning across image classification, object detection, and instance segmentation. Distillation results demonstrate that AHTD consistently outperforms its previous counterparts by a clear margin and maintains a pleasing level of methodology simplicity and training efficiency, manifesting its superiority and scalability.
KW - Difficulty Imbalance
KW - Knowledge Distillation
KW - Label Imbalance
KW - Multi-Label Learning
KW - Temperature Scaling
UR - https://www.scopus.com/pages/publications/105039641585
U2 - 10.1109/TCSVT.2026.3694911
DO - 10.1109/TCSVT.2026.3694911
M3 - 文章
AN - SCOPUS:105039641585
SN - 1051-8215
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
ER -