摘要
Remote Sensing Image-Text Retrieval (RSITR) aims to bridge the semantic gap between heterogeneous modalities and plays a vital role in various geospatial applications. As a lower dimensional and more concise modality than image, the text modality is more discriminative and may dominate the optimization process. The nonnegligible imbalanced cross-modal optimization remains a bottleneck to enhancing the model's performance. To address this issue, this study proposes a Representation Discrepancy Bridging (RDB) method for the RSITR task. On the one hand, a Cross-Modal Asymmetric Adapter (CMAA) is designed to enable modality-specific optimization and improve feature alignment. The CMAA comprises a Visual Enhancement Adapter (VEA) and a Text Semantic Adapter (TSA). VEA mines fine-grained image features through Differential Attention (DA) mechanism, while TSA identifies key textual semantics through Hierarchical Attention (HA) mechanism. On the other hand, this study extends the traditional single-task retrieval framework to a dual-task optimization framework and develops a Dual-Task Consistency Loss (DTCL). The DTCL improves cross-modal alignment robustness through an adaptive weighted combination of cross-modal, classification, and exponential moving average consistency constraints. Experiments on RSICD and RSITMD datasets show that the proposed RDB method achieves a 6 %–11 % improvement in mR metrics compared to state-of-the-art Parameter-Efficient Fine-Tuning (PEFT) methods and a 1.15 %–2 % improvement over the fully fine-tuned GeoRSCLIP model.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 130915 |
| 期刊 | Neurocomputing |
| 卷 | 650 |
| DOI | |
| 出版状态 | 已出版 - 14 10月 2025 |
| 已对外发布 | 是 |
指纹
探究 'Representation discrepancy bridging method for remote sensing image-text retrieval' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver