Abstract
Current aerial video recognition only uses vision modality to predict fixed class probabilities and does not have open-set or zero-shot recognition capabilities. We strengthen aerial video representation with linguistic semantic supervision for the first time, modeling this task as a novel vision-language multi-modality aerial video recognition paradigm. Recent studies have effectively adapted the CLIP model to natural videos. However, the diverse viewpoints, posture changes, and fast motion of UAVs introduce challenges of complex spatio-temporal contextual understanding and dynamic spatial variation capture for the aerial video-text joint representation. In this paper, we propose a contrastive learning framework with Multi-granularity Visual Prompt Co-operative information flow (MVPC-CLIP), adapting the CLIP to zero-shot, few-shot, and fully-supervised settings. 1) To absorb effective spatiotemporal clues from complex contexts, we develop frame prompts and video prompts to enhance inter-frame interaction and temporal aggregation in the visual branch. Frame prompts provide local spatiotemporal clues to enhance frame embedding and patch embedding. Video prompts with global spatiotemporal clues help frame embedding collect frame-level temporal features. 2) To capture dynamic spatial variations, we design global spatial prompts to be equipped in the adaptive spatial text mixer. The mixer can adapt to different aerial videos according to the global video spatial information, and generate more robust and discriminative textual embedding. 3) We construct novel zero-shot and few-shot settings and extensively investigate existing benchmarks. Comprehensive evaluations reveal that MVPC-CLIP achieves state-of-the-art performance across all settings. The complete codebase and few-shot dataset splits will be released.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Aerial video recognition
- multi-granularity visual prompt
- vision-language multi-modality
Fingerprint
Dive into the research topics of 'MVPC-CLIP: Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver