跳到主要导航 跳到搜索 跳到主要内容

Semantic contrastive learning via VLM for few-shot remote sensing object detection

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

Few-shot object detection (FSOD) aims to detect novel object categories with only a few annotations. However, it remains highly challenging due to limited supervision and high inter-class similarity in remote sensing images. A critical obstacle is category confusion, stemming from feature manifold entanglement caused by high inter-class visual similarity in top-down views. We identify that standard FSOD models fail to decouple discriminative cues from shared appearance features. This leads to the collapse of feature distributions within visually similar categories, severely compromising the decision boundaries. To address this problem, we propose an innovative FSOD framework based on semantic contrastive learning that leverages class-level textual knowledge from a large-scale pre-trained Vision-Language Model (VLM). Our method introduces two complementary components: (1) a Contrast-Aware Hyper-Weight (CAHW) module that generates adaptive classification weights by integrating semantic guidance, and (2) a Semantic Disambiguation Contrast (SDC) mechanism that aligns the visual features of proposals with textual prototypes to enhance inter-class separability. The integration of CAHW and SDC effectively mitigates category confusion between visually similar categories, enabling more robust and interpretable few-shot detection in remote sensing images. Extensive experiments on two standard benchmarks (DIOR and NWPU VHR-10.V2) demonstrate the effectiveness of our approach. The proposed method consistently outperforms existing state-of-the-art techniques across various k-shot settings. Code: https://github.com/Ybowei/SCL.

源语言英语
文章编号113465
期刊Pattern Recognition
178
DOI
出版状态已出版 - 10月 2026

学术指纹

探究 'Semantic contrastive learning via VLM for few-shot remote sensing object detection' 的科研主题。它们共同构成独一无二的学术指纹。

引用此