跳到主要导航 跳到搜索 跳到主要内容

MVPC-CLIP: Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

3 引用 (Scopus)

摘要

Current aerial video recognition only uses vision modality to predict fixed class probabilities and does not have open-set or zero-shot recognition capabilities. We strengthen aerial video representation with linguistic semantic supervision for the first time, modeling this task as a novel vision-language multi-modality aerial video recognition paradigm. Recent studies have effectively adapted the CLIP model to natural videos. However, the diverse viewpoints, posture changes, and fast motion of UAVs introduce challenges of complex spatio-temporal contextual understanding and dynamic spatial variation capture for the aerial video-text joint representation. In this paper, we propose a contrastive learning framework with Multi-granularity Visual Prompt Co-operative information flow (MVPC-CLIP), adapting the CLIP to zero-shot, few-shot, and fully-supervised settings. 1) To absorb effective spatiotemporal clues from complex contexts, we develop frame prompts and video prompts to enhance inter-frame interaction and temporal aggregation in the visual branch. Frame prompts provide local spatiotemporal clues to enhance frame embedding and patch embedding. Video prompts with global spatiotemporal clues help frame embedding collect frame-level temporal features. 2) To capture dynamic spatial variations, we design global spatial prompts to be equipped in the adaptive spatial text mixer. The mixer can adapt to different aerial videos according to the global video spatial information, and generate more robust and discriminative textual embedding. 3) We construct novel zero-shot and few-shot settings and extensively investigate existing benchmarks. Comprehensive evaluations reveal that MVPC-CLIP achieves state-of-the-art performance across all settings. The complete codebase and few-shot dataset splits will be released.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'MVPC-CLIP: Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此