Abstract
Open-vocabulary object detection (OVOD) has shown significant potential in real-world remote sensing applications, thanks to its adaptability to open category spaces. Existing OVOD methods mainly rely on pretrained vision-language models (VLMs) to recognize unknown categories. However, VLMs are usually trained on image-text pairs. Whilst possessing strong global semantic modeling capabilities, they struggle to accurately capture the relationship between object instances and their precise locations. To address this problem, this article presents LILK, which learns instance-level knowledge from image-level data through weak supervision. Specifically, an image-level weakly supervised knowledge injection (IWKI) module is first created, which represents instance-level semantic information by introducing global classification queries, while aligning image-level supervision with instance-level supervision through a distinct query promotion (DQP) strategy. Second, a quality-aware pseudo-label rectification (QPLR) module is developed, which filters candidate boxes using dual thresholds based on detector confidence and semantic consistency, while incorporating a SAM2-based relocation mechanism to enhance spatial localization accuracy of pseudo labels. Finally, an image-guided query enhancement (IGQE) module is introduced to guide and reinforce detection queries using instance-level image features, providing the detector with stable and modality-consistent visual References. Extensive experimental results on three commonly used remote sensing object detection datasets, DIOR, DOTA, and NWPU VHR-10, demonstrate the efficacy of LILK in performing OVOD tasks.
| Original language | English |
|---|---|
| Article number | 5626618 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 64 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Object detection
- open vocabulary
- remote sensing imagery
Fingerprint
Dive into the research topics of 'Learning Instance-Level Knowledge With Image-Level Supervision for Open-Vocabulary Object Detection in Remote Sensing Images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver