跳到主要导航 跳到搜索 跳到主要内容

Adaptively Hard-Aware Temperature Scaling for Multi-Label Distillation

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

摘要

This article aims to fill the research gap of the biased learning problem in logit knowledge distillation (KD) under the multi-label learning (MLL) context. We construct relations between temperature scaling, a critical technique in logit KD, and two MLL inherent problems: positive-negative label imbalance and learning difficulty imbalance. With this perspective, we first uncover that the existential logit KD schemes, both theoretically, lead to insufficient distillation on the beneficial positive labels and, empirically, yield poor KD performance. In particular, we introduce the tempered sigmoid and demonstrate its hard-mining effect, with the tempered teacher-student predictive discrepancy governing the penalty strength per label. Building upon these findings, we further put forth a novel KD method dubbed AHTD, dynamically balancing the distilling effort put on hard-positive labels and otherwise. Incorporating both the ground-truth and the teacher’s and student’s predictions, AHTD is instantiated by: 1) an adaptive indicator to track the hard-positive labels along the distillation course; 2) a predictive discrepancy guided scaling strategy, which is aware of the hardness level regarding specific image, label, and distillation stage. Our work is accompanied by extensive experiments on MS-COCO, PASCAL-VOC, NUS-WIDE, OpenImage, and LVIS with multiple teacher-student pairs spanning across image classification, object detection, and instance segmentation. Distillation results demonstrate that AHTD consistently outperforms its previous counterparts by a clear margin and maintains a pleasing level of methodology simplicity and training efficiency, manifesting its superiority and scalability.

指纹

探究 'Adaptively Hard-Aware Temperature Scaling for Multi-Label Distillation' 的科研主题。它们共同构成独一无二的指纹。

引用此