TY - JOUR
T1 - Dynamic Semantic Graph-Guided Cross-Modal Attention Network for Open-Vocabulary Object Detection in Aerial Images
AU - Li, Lingjun
AU - Zhang, Xinxin
AU - Feng, Xiaoxu
AU - Guo, Yanbu
AU - Liu, Shigang
AU - Peng, Yali
AU - Yao, Xiwen
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Open-vocabulary object detection (OVD) in aerial images is critical for intelligent perception of multicategory objects in complex scenarios, yet still faces two key challenges: 1) the prevalence of long-tailed categories in aerial image domains leads to insufficient visual-semantic correlation modeling for novel classes due to annotation scarcity and 2) complex backgrounds induce feature confusion and degrade the robustness of cross-modal semantic alignment. To address these issues, we propose a dynamic semantic graph-guided cross-modal attention (DynaGraph-CrossAtt) Network for aerial open-vocabulary detection. The overall architecture of DynaGraph-CrossAtt is collaboratively constituted by two core components, i.e., a dynamic semantic-aware graph (DSG) convolutional module and a bidirectional cross-modal attention (BiCMA) mechanism. In particular, DSG constructs text semantic graphs to propagate fine-grained semantic relationships (e.g., part interactions and spatial dependencies) and dynamically updates graph node representations, thereby enhancing the contextual awareness and discriminability of text embedding, effectively alleviating the representation deficiency of novel classes. Subsequently, BiCMA receives the enhanced text semantic embedding output by DSG and performs cross-modal interaction with vision features. In BiCMA, the vision-to-text branch employs dynamic weight allocation to emphasize critical semantic nodes, while the text-to-vision branch leverages semantic graphs to guide vision feature focusing on object regions, suppressing background noise, forming a dynamic soft-alignment paradigm of 'semantic-guided localization and vision feedback optimization' for cross-modal features. Experiments on aerial benchmarks demonstrate that DynaGraph-CrossAtt significantly improves OVD performance in complex aerial image scenarios.
AB - Open-vocabulary object detection (OVD) in aerial images is critical for intelligent perception of multicategory objects in complex scenarios, yet still faces two key challenges: 1) the prevalence of long-tailed categories in aerial image domains leads to insufficient visual-semantic correlation modeling for novel classes due to annotation scarcity and 2) complex backgrounds induce feature confusion and degrade the robustness of cross-modal semantic alignment. To address these issues, we propose a dynamic semantic graph-guided cross-modal attention (DynaGraph-CrossAtt) Network for aerial open-vocabulary detection. The overall architecture of DynaGraph-CrossAtt is collaboratively constituted by two core components, i.e., a dynamic semantic-aware graph (DSG) convolutional module and a bidirectional cross-modal attention (BiCMA) mechanism. In particular, DSG constructs text semantic graphs to propagate fine-grained semantic relationships (e.g., part interactions and spatial dependencies) and dynamically updates graph node representations, thereby enhancing the contextual awareness and discriminability of text embedding, effectively alleviating the representation deficiency of novel classes. Subsequently, BiCMA receives the enhanced text semantic embedding output by DSG and performs cross-modal interaction with vision features. In BiCMA, the vision-to-text branch employs dynamic weight allocation to emphasize critical semantic nodes, while the text-to-vision branch leverages semantic graphs to guide vision feature focusing on object regions, suppressing background noise, forming a dynamic soft-alignment paradigm of 'semantic-guided localization and vision feedback optimization' for cross-modal features. Experiments on aerial benchmarks demonstrate that DynaGraph-CrossAtt significantly improves OVD performance in complex aerial image scenarios.
KW - Aerial images
KW - cross-modal attention
KW - graph convolution
KW - open-vocabulary object detection (OVD)
UR - https://www.scopus.com/pages/publications/105038617127
U2 - 10.1109/TGRS.2026.3690207
DO - 10.1109/TGRS.2026.3690207
M3 - 文章
AN - SCOPUS:105038617127
SN - 0196-2892
VL - 64
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 4704012
ER -