Abstract
Typically, knowledge distillation (KD) is executed by aligning the predictions between the teacher and student with Kullback–Leibler (KL) divergence. Despite its success, this classical KD encounters two bottlenecks: (i) overlooking the rich cross-category interrelation, and (ii) enforcing a mandatory exact match of predictions. To address these issues, we present a novel order-preserving knowledge distillation (OPKD) approach. Specifically, building upon the transitive of the binary relation, OPKD aims to penalize the inconsistent of relative rank over all cross-category pairs between the teacher and student, accordingly delivering an explicit preserving of the logit order. Compared to the existing methods, OPKD leverages the cross-category information and, simultaneously, eliminates the strict requirement of actual logit value alignment. In addition, an instance-specific and pair-wise modulating factor is devised, regulating the learning effort on beneficial cross-category pairs and facilitating the KD performance. With loss and gradient analysis, we demonstrate that OPKD provides an effect of hard-mining by governing the KD process according to the difficulty in recovering relative rank. Our work is accompanied by extensive experiments on standard benchmarks across several architectures. Distillation results evidence that our OPKD consistently exceeds the prior counterparts over standard KD, self-KD, and detection KD, illustrating its superiority and versatility. Code will be made available at https://github.com/swift1988.
| Original language | English |
|---|---|
| Article number | 114405 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
Keywords
- Computer vision
- Knowledge distillation
- Model compression
Fingerprint
Dive into the research topics of 'Order-preserving knowledge distillation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver