跳到主要导航 跳到搜索 跳到主要内容

Dynamic Semantic Graph-Guided Cross-Modal Attention Network for Open-Vocabulary Object Detection in Aerial Images

  • Lingjun Li
  • , Xinxin Zhang
  • , Xiaoxu Feng
  • , Yanbu Guo
  • , Shigang Liu
  • , Yali Peng
  • , Xiwen Yao
  • Zhengzhou University of Light Industry
  • Hefei Comprehensive National Science Center
  • Shaanxi Normal University

科研成果: 期刊稿件文章同行评审

摘要

Open-vocabulary object detection (OVD) in aerial images is critical for intelligent perception of multicategory objects in complex scenarios, yet still faces two key challenges: 1) the prevalence of long-tailed categories in aerial image domains leads to insufficient visual-semantic correlation modeling for novel classes due to annotation scarcity and 2) complex backgrounds induce feature confusion and degrade the robustness of cross-modal semantic alignment. To address these issues, we propose a dynamic semantic graph-guided cross-modal attention (DynaGraph-CrossAtt) Network for aerial open-vocabulary detection. The overall architecture of DynaGraph-CrossAtt is collaboratively constituted by two core components, i.e., a dynamic semantic-aware graph (DSG) convolutional module and a bidirectional cross-modal attention (BiCMA) mechanism. In particular, DSG constructs text semantic graphs to propagate fine-grained semantic relationships (e.g., part interactions and spatial dependencies) and dynamically updates graph node representations, thereby enhancing the contextual awareness and discriminability of text embedding, effectively alleviating the representation deficiency of novel classes. Subsequently, BiCMA receives the enhanced text semantic embedding output by DSG and performs cross-modal interaction with vision features. In BiCMA, the vision-to-text branch employs dynamic weight allocation to emphasize critical semantic nodes, while the text-to-vision branch leverages semantic graphs to guide vision feature focusing on object regions, suppressing background noise, forming a dynamic soft-alignment paradigm of 'semantic-guided localization and vision feedback optimization' for cross-modal features. Experiments on aerial benchmarks demonstrate that DynaGraph-CrossAtt significantly improves OVD performance in complex aerial image scenarios.

源语言英语
文章编号4704012
期刊IEEE Transactions on Geoscience and Remote Sensing
64
DOI
出版状态已出版 - 2026

指纹

探究 'Dynamic Semantic Graph-Guided Cross-Modal Attention Network for Open-Vocabulary Object Detection in Aerial Images' 的科研主题。它们共同构成独一无二的指纹。

引用此