TY - JOUR
T1 - Order-preserving knowledge distillation
AU - Li, Cong
AU - Cheng, Gong
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/12
Y1 - 2026/12
N2 - Typically, knowledge distillation (KD) is executed by aligning the predictions between the teacher and student with Kullback–Leibler (KL) divergence. Despite its success, this classical KD encounters two bottlenecks: (i) overlooking the rich cross-category interrelation, and (ii) enforcing a mandatory exact match of predictions. To address these issues, we present a novel order-preserving knowledge distillation (OPKD) approach. Specifically, building upon the transitive of the binary relation, OPKD aims to penalize the inconsistent of relative rank over all cross-category pairs between the teacher and student, accordingly delivering an explicit preserving of the logit order. Compared to the existing methods, OPKD leverages the cross-category information and, simultaneously, eliminates the strict requirement of actual logit value alignment. In addition, an instance-specific and pair-wise modulating factor is devised, regulating the learning effort on beneficial cross-category pairs and facilitating the KD performance. With loss and gradient analysis, we demonstrate that OPKD provides an effect of hard-mining by governing the KD process according to the difficulty in recovering relative rank. Our work is accompanied by extensive experiments on standard benchmarks across several architectures. Distillation results evidence that our OPKD consistently exceeds the prior counterparts over standard KD, self-KD, and detection KD, illustrating its superiority and versatility. Code will be made available at https://github.com/swift1988.
AB - Typically, knowledge distillation (KD) is executed by aligning the predictions between the teacher and student with Kullback–Leibler (KL) divergence. Despite its success, this classical KD encounters two bottlenecks: (i) overlooking the rich cross-category interrelation, and (ii) enforcing a mandatory exact match of predictions. To address these issues, we present a novel order-preserving knowledge distillation (OPKD) approach. Specifically, building upon the transitive of the binary relation, OPKD aims to penalize the inconsistent of relative rank over all cross-category pairs between the teacher and student, accordingly delivering an explicit preserving of the logit order. Compared to the existing methods, OPKD leverages the cross-category information and, simultaneously, eliminates the strict requirement of actual logit value alignment. In addition, an instance-specific and pair-wise modulating factor is devised, regulating the learning effort on beneficial cross-category pairs and facilitating the KD performance. With loss and gradient analysis, we demonstrate that OPKD provides an effect of hard-mining by governing the KD process according to the difficulty in recovering relative rank. Our work is accompanied by extensive experiments on standard benchmarks across several architectures. Distillation results evidence that our OPKD consistently exceeds the prior counterparts over standard KD, self-KD, and detection KD, illustrating its superiority and versatility. Code will be made available at https://github.com/swift1988.
KW - Computer vision
KW - Knowledge distillation
KW - Model compression
UR - https://www.scopus.com/pages/publications/105044811284
U2 - 10.1016/j.patcog.2026.114405
DO - 10.1016/j.patcog.2026.114405
M3 - 文章
AN - SCOPUS:105044811284
SN - 0031-3203
VL - 180
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 114405
ER -