Abstract
Robust visual tracking requires a stable spatio-temporal representation of the target state across video frames. Existing spatio-temporal trackers typically pass historical tokens or target states forward as auxiliary cues, but they do not directly model how these states evolve in feature space. Temporal information is therefore used mainly from one frame to the next, which can weaken tracking under occlusion, drift, and long-term appearance changes. To address this limitation, we formulate target-state evolution as a continuous-time dynamical system and propose FIST, a flow-inspired framework for spatio-temporal target-state modeling. Motivated by Flow Matching, FIST learns a vector field over the latent target-state distribution and uses it to estimate both historical and future target states. This provides an explicit description of temporal state evolution. FIST further combines cross-frame temporal state modeling with cross-layer spatial state compression, producing a compact spatio-temporal representation that retains temporal continuity and multi-level spatial information for target localization. Experiments on seven tracking benchmarks show that FISTrack balances accuracy and efficiency. FISTrack-B384 reaches 80.0% AO on GOT-10k, while FISTrack-T256 runs at 196 FPS on GPU and 35 FPS on CPU with only 8.78M parameters. Notably, FIST can be integrated into different transformer-based trackers. The code and trained models will be released at https://github.com/ElliottZhen/FIST.
| Original language | English |
|---|---|
| Article number | 114243 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
Keywords
- Flow matching
- Generality
- Spatio-temporal target state
- Target state modeling
- Visual object tracking
Fingerprint
Dive into the research topics of 'FIST: Flow-inspired spatio-temporal target state modeling for visual tracking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver