Abstract
Effectively modeling spatial-temporal consistent target dependency is essential for robust visual tracking. Existing solutions — template updating and token propagation mechanisms — suffer from non-trivial hyperparameter tuning that compromises portability and generalizability. To address this, we propose Mamba in Attention (MiA), a hyperparameter-free plug-and-play module that mitigates the need for manual tuning while effectively enhancing spatial-temporal consistency. The MiA module comprises two complementary branches: an Attention-based spatial branch that captures static spatial relationships, and a Mamba-based temporal branch, which continuously encodes historical context via hidden states of the State Space Model (SSM). By fusing these two dependencies, MiA enables static-dynamic interaction, significantly improving the robustness of target localization. For extreme scenarios, we further contribute a Channel-aware KAN (CKAN), first introducing Kolmogorov-Arnold Networks to visual tracking for enhanced nonlinear channel relationship representation. Built upon MiA and CKAN, we present a state-of-the-art lightweight tracker, MiATrack, which achieves an excellent performance-speed-parameter trade-off on all five commonly used tracking benchmarks. Notably, MiA offers seamless integration into existing tracking frameworks, delivering performance gains with only 30 training epochs, minimal overhead, and zero hyperparameter tuning. Extensive experiments across five diverse trackers and six tracking benchmarks validate the excellent portability and generalizability of our proposed MiA module. Code is available at https://github.com/Xiaochen918/MiATrack.
| Original language | English |
|---|---|
| Article number | 114033 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
Keywords
- KAN
- Mamba
- Spatial–temporal modeling
- Visual object tracking
Fingerprint
Dive into the research topics of 'MiA: A plug-and-play hyperparameter-free Mamba in Attention module for spatial–temporal consistent visual tracking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver