Abstract
This paper presents a novel video-level RGB-T tracking paradigm based on prompt learning, termed PromptTrack, which establishes dense spatial-temporal associations through cross-modal interactions. The method introduces streaming temporal prompts to capture continuous target dynamics (e.g., appearance changes and motion trajectories), while leveraging multimodal spatial prompts to utilize complementary RGB and thermal-infrared (TIR) features dynamically. By propagating temporal prompts through consecutive frames and integrating bidirectional spatial interactions between modalities, PromptTrack achieves superior tracking performance in complex scenarios such as occlusion, low illumination, and distractors. The proposed framework employs a unified multimodal encoder with spatial-temporal modeling via multimodal spatial prompt blocks, enabling efficient fusion of RGB-TIR features without requiring domain-specific structure modifications. Extensive experiments on three RGB-T benchmarks (LasHeR, RGBT210, RGBT234) demonstrate that PromptTrack achieves new state-of-the-art performance, with 76.2% in precision rate on LasHeR and outperforming existing methods by +1.9% in precision rate and +0.5% in success rate on RGBT234. Notably, its modality-agnostic design facilitates seamless generalization to RGB-D and RGB-E tracking domains, achieving new benchmarks on DepthTrack, VOT-RGBD2022, and VisEvent datasets.
| Original language | English |
|---|---|
| Pages (from-to) | 5175-5189 |
| Number of pages | 15 |
| Journal | IEEE Transactions on Multimedia |
| Volume | 28 |
| DOIs | |
| State | Published - 2026 |
Keywords
- RGB-T tracking
- multimodal
- prompt learning
- spatial-temporal modeling
Fingerprint
Dive into the research topics of 'PromptTrack: Streaming Spatial-Temporal Prompt Learning for RGB-T Tracking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver