TY - JOUR
T1 - MVPC-CLIP
T2 - Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition
AU - Zhan, Yang
AU - Yuan, Yuan
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Current aerial video recognition only uses vision modality to predict fixed class probabilities and does not have open-set or zero-shot recognition capabilities. We strengthen aerial video representation with linguistic semantic supervision for the first time, modeling this task as a novel vision-language multi-modality aerial video recognition paradigm. Recent studies have effectively adapted the CLIP model to natural videos. However, the diverse viewpoints, posture changes, and fast motion of UAVs introduce challenges of complex spatio-temporal contextual understanding and dynamic spatial variation capture for the aerial video-text joint representation. In this paper, we propose a contrastive learning framework with Multi-granularity Visual Prompt Co-operative information flow (MVPC-CLIP), adapting the CLIP to zero-shot, few-shot, and fully-supervised settings. 1) To absorb effective spatiotemporal clues from complex contexts, we develop frame prompts and video prompts to enhance inter-frame interaction and temporal aggregation in the visual branch. Frame prompts provide local spatiotemporal clues to enhance frame embedding and patch embedding. Video prompts with global spatiotemporal clues help frame embedding collect frame-level temporal features. 2) To capture dynamic spatial variations, we design global spatial prompts to be equipped in the adaptive spatial text mixer. The mixer can adapt to different aerial videos according to the global video spatial information, and generate more robust and discriminative textual embedding. 3) We construct novel zero-shot and few-shot settings and extensively investigate existing benchmarks. Comprehensive evaluations reveal that MVPC-CLIP achieves state-of-the-art performance across all settings. The complete codebase and few-shot dataset splits will be released.
AB - Current aerial video recognition only uses vision modality to predict fixed class probabilities and does not have open-set or zero-shot recognition capabilities. We strengthen aerial video representation with linguistic semantic supervision for the first time, modeling this task as a novel vision-language multi-modality aerial video recognition paradigm. Recent studies have effectively adapted the CLIP model to natural videos. However, the diverse viewpoints, posture changes, and fast motion of UAVs introduce challenges of complex spatio-temporal contextual understanding and dynamic spatial variation capture for the aerial video-text joint representation. In this paper, we propose a contrastive learning framework with Multi-granularity Visual Prompt Co-operative information flow (MVPC-CLIP), adapting the CLIP to zero-shot, few-shot, and fully-supervised settings. 1) To absorb effective spatiotemporal clues from complex contexts, we develop frame prompts and video prompts to enhance inter-frame interaction and temporal aggregation in the visual branch. Frame prompts provide local spatiotemporal clues to enhance frame embedding and patch embedding. Video prompts with global spatiotemporal clues help frame embedding collect frame-level temporal features. 2) To capture dynamic spatial variations, we design global spatial prompts to be equipped in the adaptive spatial text mixer. The mixer can adapt to different aerial videos according to the global video spatial information, and generate more robust and discriminative textual embedding. 3) We construct novel zero-shot and few-shot settings and extensively investigate existing benchmarks. Comprehensive evaluations reveal that MVPC-CLIP achieves state-of-the-art performance across all settings. The complete codebase and few-shot dataset splits will be released.
KW - Aerial video recognition
KW - multi-granularity visual prompt
KW - vision-language multi-modality
UR - https://www.scopus.com/pages/publications/105032150515
U2 - 10.1109/TMM.2026.3668495
DO - 10.1109/TMM.2026.3668495
M3 - 文章
AN - SCOPUS:105032150515
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -