摘要
Few-shot object detection (FSOD) aims to detect novel object categories with only a few annotations. However, it remains highly challenging due to limited supervision and high inter-class similarity in remote sensing images. A critical obstacle is category confusion, stemming from feature manifold entanglement caused by high inter-class visual similarity in top-down views. We identify that standard FSOD models fail to decouple discriminative cues from shared appearance features. This leads to the collapse of feature distributions within visually similar categories, severely compromising the decision boundaries. To address this problem, we propose an innovative FSOD framework based on semantic contrastive learning that leverages class-level textual knowledge from a large-scale pre-trained Vision-Language Model (VLM). Our method introduces two complementary components: (1) a Contrast-Aware Hyper-Weight (CAHW) module that generates adaptive classification weights by integrating semantic guidance, and (2) a Semantic Disambiguation Contrast (SDC) mechanism that aligns the visual features of proposals with textual prototypes to enhance inter-class separability. The integration of CAHW and SDC effectively mitigates category confusion between visually similar categories, enabling more robust and interpretable few-shot detection in remote sensing images. Extensive experiments on two standard benchmarks (DIOR and NWPU VHR-10.V2) demonstrate the effectiveness of our approach. The proposed method consistently outperforms existing state-of-the-art techniques across various k-shot settings. Code: https://github.com/Ybowei/SCL.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 113465 |
| 期刊 | Pattern Recognition |
| 卷 | 178 |
| DOI | |
| 出版状态 | 已出版 - 10月 2026 |
学术指纹
探究 'Semantic contrastive learning via VLM for few-shot remote sensing object detection' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver