TY - JOUR
T1 - Learning Instance-Level Knowledge With Image-Level Supervision for Open-Vocabulary Object Detection in Remote Sensing Images
AU - Li, Yan
AU - Bai, Yunpeng
AU - Ma, Jiaman
AU - Temirbayev, Amirkhan
AU - Li, Ying
AU - Shang, Changjing
AU - Shen, Qiang
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Open-vocabulary object detection (OVOD) has shown significant potential in real-world remote sensing applications, thanks to its adaptability to open category spaces. Existing OVOD methods mainly rely on pretrained vision-language models (VLMs) to recognize unknown categories. However, VLMs are usually trained on image-text pairs. Whilst possessing strong global semantic modeling capabilities, they struggle to accurately capture the relationship between object instances and their precise locations. To address this problem, this article presents LILK, which learns instance-level knowledge from image-level data through weak supervision. Specifically, an image-level weakly supervised knowledge injection (IWKI) module is first created, which represents instance-level semantic information by introducing global classification queries, while aligning image-level supervision with instance-level supervision through a distinct query promotion (DQP) strategy. Second, a quality-aware pseudo-label rectification (QPLR) module is developed, which filters candidate boxes using dual thresholds based on detector confidence and semantic consistency, while incorporating a SAM2-based relocation mechanism to enhance spatial localization accuracy of pseudo labels. Finally, an image-guided query enhancement (IGQE) module is introduced to guide and reinforce detection queries using instance-level image features, providing the detector with stable and modality-consistent visual References. Extensive experimental results on three commonly used remote sensing object detection datasets, DIOR, DOTA, and NWPU VHR-10, demonstrate the efficacy of LILK in performing OVOD tasks.
AB - Open-vocabulary object detection (OVOD) has shown significant potential in real-world remote sensing applications, thanks to its adaptability to open category spaces. Existing OVOD methods mainly rely on pretrained vision-language models (VLMs) to recognize unknown categories. However, VLMs are usually trained on image-text pairs. Whilst possessing strong global semantic modeling capabilities, they struggle to accurately capture the relationship between object instances and their precise locations. To address this problem, this article presents LILK, which learns instance-level knowledge from image-level data through weak supervision. Specifically, an image-level weakly supervised knowledge injection (IWKI) module is first created, which represents instance-level semantic information by introducing global classification queries, while aligning image-level supervision with instance-level supervision through a distinct query promotion (DQP) strategy. Second, a quality-aware pseudo-label rectification (QPLR) module is developed, which filters candidate boxes using dual thresholds based on detector confidence and semantic consistency, while incorporating a SAM2-based relocation mechanism to enhance spatial localization accuracy of pseudo labels. Finally, an image-guided query enhancement (IGQE) module is introduced to guide and reinforce detection queries using instance-level image features, providing the detector with stable and modality-consistent visual References. Extensive experimental results on three commonly used remote sensing object detection datasets, DIOR, DOTA, and NWPU VHR-10, demonstrate the efficacy of LILK in performing OVOD tasks.
KW - Object detection
KW - open vocabulary
KW - remote sensing imagery
UR - https://www.scopus.com/pages/publications/105041002007
U2 - 10.1109/TGRS.2026.3699883
DO - 10.1109/TGRS.2026.3699883
M3 - 文章
AN - SCOPUS:105041002007
SN - 0196-2892
VL - 64
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 5626618
ER -