Abstract
Open-vocabulary object detection (OVD) in aerial images is critical for intelligent perception of multicategory objects in complex scenarios, yet still faces two key challenges: 1) the prevalence of long-tailed categories in aerial image domains leads to insufficient visual-semantic correlation modeling for novel classes due to annotation scarcity and 2) complex backgrounds induce feature confusion and degrade the robustness of cross-modal semantic alignment. To address these issues, we propose a dynamic semantic graph-guided cross-modal attention (DynaGraph-CrossAtt) Network for aerial open-vocabulary detection. The overall architecture of DynaGraph-CrossAtt is collaboratively constituted by two core components, i.e., a dynamic semantic-aware graph (DSG) convolutional module and a bidirectional cross-modal attention (BiCMA) mechanism. In particular, DSG constructs text semantic graphs to propagate fine-grained semantic relationships (e.g., part interactions and spatial dependencies) and dynamically updates graph node representations, thereby enhancing the contextual awareness and discriminability of text embedding, effectively alleviating the representation deficiency of novel classes. Subsequently, BiCMA receives the enhanced text semantic embedding output by DSG and performs cross-modal interaction with vision features. In BiCMA, the vision-to-text branch employs dynamic weight allocation to emphasize critical semantic nodes, while the text-to-vision branch leverages semantic graphs to guide vision feature focusing on object regions, suppressing background noise, forming a dynamic soft-alignment paradigm of 'semantic-guided localization and vision feedback optimization' for cross-modal features. Experiments on aerial benchmarks demonstrate that DynaGraph-CrossAtt significantly improves OVD performance in complex aerial image scenarios.
| Original language | English |
|---|---|
| Article number | 4704012 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 64 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Aerial images
- cross-modal attention
- graph convolution
- open-vocabulary object detection (OVD)
Fingerprint
Dive into the research topics of 'Dynamic Semantic Graph-Guided Cross-Modal Attention Network for Open-Vocabulary Object Detection in Aerial Images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver