TY - JOUR
T1 - MiA
T2 - A plug-and-play hyperparameter-free Mamba in Attention module for spatial–temporal consistent visual tracking
AU - Chen, Yao
AU - Jia, Guancheng
AU - Ma, Ding
AU - Zha, Yufei
AU - Zhang, Peng
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/12
Y1 - 2026/12
N2 - Effectively modeling spatial-temporal consistent target dependency is essential for robust visual tracking. Existing solutions — template updating and token propagation mechanisms — suffer from non-trivial hyperparameter tuning that compromises portability and generalizability. To address this, we propose Mamba in Attention (MiA), a hyperparameter-free plug-and-play module that mitigates the need for manual tuning while effectively enhancing spatial-temporal consistency. The MiA module comprises two complementary branches: an Attention-based spatial branch that captures static spatial relationships, and a Mamba-based temporal branch, which continuously encodes historical context via hidden states of the State Space Model (SSM). By fusing these two dependencies, MiA enables static-dynamic interaction, significantly improving the robustness of target localization. For extreme scenarios, we further contribute a Channel-aware KAN (CKAN), first introducing Kolmogorov-Arnold Networks to visual tracking for enhanced nonlinear channel relationship representation. Built upon MiA and CKAN, we present a state-of-the-art lightweight tracker, MiATrack, which achieves an excellent performance-speed-parameter trade-off on all five commonly used tracking benchmarks. Notably, MiA offers seamless integration into existing tracking frameworks, delivering performance gains with only 30 training epochs, minimal overhead, and zero hyperparameter tuning. Extensive experiments across five diverse trackers and six tracking benchmarks validate the excellent portability and generalizability of our proposed MiA module. Code is available at https://github.com/Xiaochen918/MiATrack.
AB - Effectively modeling spatial-temporal consistent target dependency is essential for robust visual tracking. Existing solutions — template updating and token propagation mechanisms — suffer from non-trivial hyperparameter tuning that compromises portability and generalizability. To address this, we propose Mamba in Attention (MiA), a hyperparameter-free plug-and-play module that mitigates the need for manual tuning while effectively enhancing spatial-temporal consistency. The MiA module comprises two complementary branches: an Attention-based spatial branch that captures static spatial relationships, and a Mamba-based temporal branch, which continuously encodes historical context via hidden states of the State Space Model (SSM). By fusing these two dependencies, MiA enables static-dynamic interaction, significantly improving the robustness of target localization. For extreme scenarios, we further contribute a Channel-aware KAN (CKAN), first introducing Kolmogorov-Arnold Networks to visual tracking for enhanced nonlinear channel relationship representation. Built upon MiA and CKAN, we present a state-of-the-art lightweight tracker, MiATrack, which achieves an excellent performance-speed-parameter trade-off on all five commonly used tracking benchmarks. Notably, MiA offers seamless integration into existing tracking frameworks, delivering performance gains with only 30 training epochs, minimal overhead, and zero hyperparameter tuning. Extensive experiments across five diverse trackers and six tracking benchmarks validate the excellent portability and generalizability of our proposed MiA module. Code is available at https://github.com/Xiaochen918/MiATrack.
KW - KAN
KW - Mamba
KW - Spatial–temporal modeling
KW - Visual object tracking
UR - https://www.scopus.com/pages/publications/105040371102
U2 - 10.1016/j.patcog.2026.114033
DO - 10.1016/j.patcog.2026.114033
M3 - 文章
AN - SCOPUS:105040371102
SN - 0031-3203
VL - 180
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 114033
ER -