Skip to main navigation Skip to search Skip to main content

Dynamic Semantic Graph-Guided Cross-Modal Attention Network for Open-Vocabulary Object Detection in Aerial Images

  • Lingjun Li
  • , Xinxin Zhang
  • , Xiaoxu Feng
  • , Yanbu Guo
  • , Shigang Liu
  • , Yali Peng
  • , Xiwen Yao
  • Zhengzhou University of Light Industry
  • Hefei Comprehensive National Science Center
  • Shaanxi Normal University

Research output: Contribution to journalArticlepeer-review

Abstract

Open-vocabulary object detection (OVD) in aerial images is critical for intelligent perception of multicategory objects in complex scenarios, yet still faces two key challenges: 1) the prevalence of long-tailed categories in aerial image domains leads to insufficient visual-semantic correlation modeling for novel classes due to annotation scarcity and 2) complex backgrounds induce feature confusion and degrade the robustness of cross-modal semantic alignment. To address these issues, we propose a dynamic semantic graph-guided cross-modal attention (DynaGraph-CrossAtt) Network for aerial open-vocabulary detection. The overall architecture of DynaGraph-CrossAtt is collaboratively constituted by two core components, i.e., a dynamic semantic-aware graph (DSG) convolutional module and a bidirectional cross-modal attention (BiCMA) mechanism. In particular, DSG constructs text semantic graphs to propagate fine-grained semantic relationships (e.g., part interactions and spatial dependencies) and dynamically updates graph node representations, thereby enhancing the contextual awareness and discriminability of text embedding, effectively alleviating the representation deficiency of novel classes. Subsequently, BiCMA receives the enhanced text semantic embedding output by DSG and performs cross-modal interaction with vision features. In BiCMA, the vision-to-text branch employs dynamic weight allocation to emphasize critical semantic nodes, while the text-to-vision branch leverages semantic graphs to guide vision feature focusing on object regions, suppressing background noise, forming a dynamic soft-alignment paradigm of 'semantic-guided localization and vision feedback optimization' for cross-modal features. Experiments on aerial benchmarks demonstrate that DynaGraph-CrossAtt significantly improves OVD performance in complex aerial image scenarios.

Original languageEnglish
Article number4704012
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume64
DOIs
StatePublished - 2026

Keywords

  • Aerial images
  • cross-modal attention
  • graph convolution
  • open-vocabulary object detection (OVD)

Fingerprint

Dive into the research topics of 'Dynamic Semantic Graph-Guided Cross-Modal Attention Network for Open-Vocabulary Object Detection in Aerial Images'. Together they form a unique fingerprint.

Cite this