TY - JOUR
T1 - DecoderTracker
T2 - Decoder-only end-to-end method for multiple-object tracking
AU - Liao, Pan
AU - Yang, Feng
AU - Wu, Di
AU - Zhao, Wenhui
AU - Yu, Jinwen
AU - Zhang, Dingwen
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/9
Y1 - 2026/9
N2 - Decoder-only Transformer architectures, such as GPT, have demonstrated superior performance in many areas compared to traditional encoder-decoder structure transformer methods. Over the years, end-to-end methods based on the traditional transformer structure, like MOTR, have achieved remarkable performance in multi-object tracking. However, these methods suffer from substantial computational costs and optimization challenges inherent to dynamic data processing, leading to suboptimal inference speeds and prolonged training times. To address the aforementioned issues, this paper optimized the network architecture and proposed an effective training strategy to mitigate the problem of prolonged training times, thereby developing DecoderTracker, a novel end-to-end tracking method. Subsequently, to tackle the optimization challenges arising from dynamic data, this paper introduced FixDT by incorporating a Fixed-Size Query Memory and refining certain attention layers. Our methods, outperforms MOTR on multiple benchmarks without incorporating complex heuristic components, featuring a 2 to 3 times faster inference than MOTR, respectively. The proposed method is implemented in open-source code, accessible at https://github.com/liaopan-lp/MO-YOLO.
AB - Decoder-only Transformer architectures, such as GPT, have demonstrated superior performance in many areas compared to traditional encoder-decoder structure transformer methods. Over the years, end-to-end methods based on the traditional transformer structure, like MOTR, have achieved remarkable performance in multi-object tracking. However, these methods suffer from substantial computational costs and optimization challenges inherent to dynamic data processing, leading to suboptimal inference speeds and prolonged training times. To address the aforementioned issues, this paper optimized the network architecture and proposed an effective training strategy to mitigate the problem of prolonged training times, thereby developing DecoderTracker, a novel end-to-end tracking method. Subsequently, to tackle the optimization challenges arising from dynamic data, this paper introduced FixDT by incorporating a Fixed-Size Query Memory and refining certain attention layers. Our methods, outperforms MOTR on multiple benchmarks without incorporating complex heuristic components, featuring a 2 to 3 times faster inference than MOTR, respectively. The proposed method is implemented in open-source code, accessible at https://github.com/liaopan-lp/MO-YOLO.
KW - Decoder
KW - End-to-end
KW - Multi-object tracking
UR - https://www.scopus.com/pages/publications/105030180735
U2 - 10.1016/j.patcog.2026.113242
DO - 10.1016/j.patcog.2026.113242
M3 - 文章
AN - SCOPUS:105030180735
SN - 0031-3203
VL - 177
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 113242
ER -