Abstract
End-to-end multi-object tracking (MOT) methods often suffer from imbalanced training between detection and tracking queries, primarily caused by the one-by-one label assignment mechanism. A common solution involves either introducing an additional detection decoder or applying detection-specific supervision. However, the former substantially increases computational overhead, while the latter may result in unstable optimization due to implicit matching competition between detection and tracking queries. In this paper, we present HT-MOTR, a real-time end-to-end MOT framework that achieves a better balance between accuracy and efficiency. HT-MOTR introduces two key components: Proposal-Guided Hierarchical Matching (PHM) and Tracking-Conditioned Guided Decoding (TCGD). In PHM, we first generate initial detection queries based on the output proposals of an efficient hybrid encoder. Then, we perform progressive matching for detection and tracking queries at different decoder layers: early detection-prioritized matching, middle joint matching, and final tracking-prioritized matching. The TCGD module further enhances detection queries with temporal information derived from tracking queries. Specifically, tracking queries first attend to encoder features via cross-attention, and the resulting temporally enriched queries subsequently guide detection queries through self-attention before interacting with visual features. Extensive experiments demonstrate the effectiveness of our approach. HT-MOTR achieves 66.5% HOTA on DanceTrack and runs at a real-time speed of 26.7 FPS, thus validating both its robustness and efficiency. Our code and pre-trained models will be publicly released.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- End-to-end multi-object tracking
- Hierarchical matching
- Real-time performance
- Temporal modeling
Fingerprint
Dive into the research topics of 'HT-MOTR: End-to-End Multi-Object Tracking with Hierarchical Matching and Temporal Guidance'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver