TY - JOUR
T1 - FIST
T2 - Flow-inspired spatio-temporal target state modeling for visual tracking
AU - Wan, Zhen
AU - Sun, Boxiong
AU - Han, Tonghao
AU - Zha, Yufei
AU - Zhang, Peng
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/12
Y1 - 2026/12
N2 - Robust visual tracking requires a stable spatio-temporal representation of the target state across video frames. Existing spatio-temporal trackers typically pass historical tokens or target states forward as auxiliary cues, but they do not directly model how these states evolve in feature space. Temporal information is therefore used mainly from one frame to the next, which can weaken tracking under occlusion, drift, and long-term appearance changes. To address this limitation, we formulate target-state evolution as a continuous-time dynamical system and propose FIST, a flow-inspired framework for spatio-temporal target-state modeling. Motivated by Flow Matching, FIST learns a vector field over the latent target-state distribution and uses it to estimate both historical and future target states. This provides an explicit description of temporal state evolution. FIST further combines cross-frame temporal state modeling with cross-layer spatial state compression, producing a compact spatio-temporal representation that retains temporal continuity and multi-level spatial information for target localization. Experiments on seven tracking benchmarks show that FISTrack balances accuracy and efficiency. FISTrack-B384 reaches 80.0% AO on GOT-10k, while FISTrack-T256 runs at 196 FPS on GPU and 35 FPS on CPU with only 8.78M parameters. Notably, FIST can be integrated into different transformer-based trackers. The code and trained models will be released at https://github.com/ElliottZhen/FIST.
AB - Robust visual tracking requires a stable spatio-temporal representation of the target state across video frames. Existing spatio-temporal trackers typically pass historical tokens or target states forward as auxiliary cues, but they do not directly model how these states evolve in feature space. Temporal information is therefore used mainly from one frame to the next, which can weaken tracking under occlusion, drift, and long-term appearance changes. To address this limitation, we formulate target-state evolution as a continuous-time dynamical system and propose FIST, a flow-inspired framework for spatio-temporal target-state modeling. Motivated by Flow Matching, FIST learns a vector field over the latent target-state distribution and uses it to estimate both historical and future target states. This provides an explicit description of temporal state evolution. FIST further combines cross-frame temporal state modeling with cross-layer spatial state compression, producing a compact spatio-temporal representation that retains temporal continuity and multi-level spatial information for target localization. Experiments on seven tracking benchmarks show that FISTrack balances accuracy and efficiency. FISTrack-B384 reaches 80.0% AO on GOT-10k, while FISTrack-T256 runs at 196 FPS on GPU and 35 FPS on CPU with only 8.78M parameters. Notably, FIST can be integrated into different transformer-based trackers. The code and trained models will be released at https://github.com/ElliottZhen/FIST.
KW - Flow matching
KW - Generality
KW - Spatio-temporal target state
KW - Target state modeling
KW - Visual object tracking
UR - https://www.scopus.com/pages/publications/105042266159
U2 - 10.1016/j.patcog.2026.114243
DO - 10.1016/j.patcog.2026.114243
M3 - 文章
AN - SCOPUS:105042266159
SN - 0031-3203
VL - 180
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 114243
ER -