TY - JOUR
T1 - HT-MOTR
T2 - End-to-End Multi-Object Tracking with Hierarchical Matching and Temporal Guidance
AU - Li, Yuhao
AU - Sun, Jinqiu
AU - Zhu, Yu
AU - Zhang, Yanning
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - End-to-end multi-object tracking (MOT) methods often suffer from imbalanced training between detection and tracking queries, primarily caused by the one-by-one label assignment mechanism. A common solution involves either introducing an additional detection decoder or applying detection-specific supervision. However, the former substantially increases computational overhead, while the latter may result in unstable optimization due to implicit matching competition between detection and tracking queries. In this paper, we present HT-MOTR, a real-time end-to-end MOT framework that achieves a better balance between accuracy and efficiency. HT-MOTR introduces two key components: Proposal-Guided Hierarchical Matching (PHM) and Tracking-Conditioned Guided Decoding (TCGD). In PHM, we first generate initial detection queries based on the output proposals of an efficient hybrid encoder. Then, we perform progressive matching for detection and tracking queries at different decoder layers: early detection-prioritized matching, middle joint matching, and final tracking-prioritized matching. The TCGD module further enhances detection queries with temporal information derived from tracking queries. Specifically, tracking queries first attend to encoder features via cross-attention, and the resulting temporally enriched queries subsequently guide detection queries through self-attention before interacting with visual features. Extensive experiments demonstrate the effectiveness of our approach. HT-MOTR achieves 66.5% HOTA on DanceTrack and runs at a real-time speed of 26.7 FPS, thus validating both its robustness and efficiency. Our code and pre-trained models will be publicly released.
AB - End-to-end multi-object tracking (MOT) methods often suffer from imbalanced training between detection and tracking queries, primarily caused by the one-by-one label assignment mechanism. A common solution involves either introducing an additional detection decoder or applying detection-specific supervision. However, the former substantially increases computational overhead, while the latter may result in unstable optimization due to implicit matching competition between detection and tracking queries. In this paper, we present HT-MOTR, a real-time end-to-end MOT framework that achieves a better balance between accuracy and efficiency. HT-MOTR introduces two key components: Proposal-Guided Hierarchical Matching (PHM) and Tracking-Conditioned Guided Decoding (TCGD). In PHM, we first generate initial detection queries based on the output proposals of an efficient hybrid encoder. Then, we perform progressive matching for detection and tracking queries at different decoder layers: early detection-prioritized matching, middle joint matching, and final tracking-prioritized matching. The TCGD module further enhances detection queries with temporal information derived from tracking queries. Specifically, tracking queries first attend to encoder features via cross-attention, and the resulting temporally enriched queries subsequently guide detection queries through self-attention before interacting with visual features. Extensive experiments demonstrate the effectiveness of our approach. HT-MOTR achieves 66.5% HOTA on DanceTrack and runs at a real-time speed of 26.7 FPS, thus validating both its robustness and efficiency. Our code and pre-trained models will be publicly released.
KW - End-to-end multi-object tracking
KW - Hierarchical matching
KW - Real-time performance
KW - Temporal modeling
UR - https://www.scopus.com/pages/publications/105043407377
U2 - 10.1109/TMM.2026.3705893
DO - 10.1109/TMM.2026.3705893
M3 - 文章
AN - SCOPUS:105043407377
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -