跳到主要导航 跳到搜索 跳到主要内容

FIST: Flow-inspired spatio-temporal target state modeling for visual tracking

  • Zhen Wan
  • , Boxiong Sun
  • , Tonghao Han
  • , Yufei Zha
  • , Peng Zhang
  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

摘要

Robust visual tracking requires a stable spatio-temporal representation of the target state across video frames. Existing spatio-temporal trackers typically pass historical tokens or target states forward as auxiliary cues, but they do not directly model how these states evolve in feature space. Temporal information is therefore used mainly from one frame to the next, which can weaken tracking under occlusion, drift, and long-term appearance changes. To address this limitation, we formulate target-state evolution as a continuous-time dynamical system and propose FIST, a flow-inspired framework for spatio-temporal target-state modeling. Motivated by Flow Matching, FIST learns a vector field over the latent target-state distribution and uses it to estimate both historical and future target states. This provides an explicit description of temporal state evolution. FIST further combines cross-frame temporal state modeling with cross-layer spatial state compression, producing a compact spatio-temporal representation that retains temporal continuity and multi-level spatial information for target localization. Experiments on seven tracking benchmarks show that FISTrack balances accuracy and efficiency. FISTrack-B384 reaches 80.0% AO on GOT-10k, while FISTrack-T256 runs at 196 FPS on GPU and 35 FPS on CPU with only 8.78M parameters. Notably, FIST can be integrated into different transformer-based trackers. The code and trained models will be released at https://github.com/ElliottZhen/FIST.

源语言英语
文章编号114243
期刊Pattern Recognition
180
DOI
出版状态已出版 - 12月 2026

指纹

探究 'FIST: Flow-inspired spatio-temporal target state modeling for visual tracking' 的科研主题。它们共同构成独一无二的指纹。

引用此