跳到主要导航 跳到搜索 跳到主要内容

Representation discrepancy bridging method for remote sensing image-text retrieval

  • Hailong Ning
  • , Siying Wang
  • , Tao Lei
  • , Xiaopeng Cao
  • , Huanmin Dou
  • , Bin Zhao
  • , Asoke K. Nandi
  • , Petia Radeva
  • Xi'an Institute of Posts and Telecommunications
  • Shaanxi Key Laboratory of Network Data Analysis and Intelligent Processing
  • Xi’an Key Laboratory of Big Data and Intelligent Computing
  • Shaanxi University of Science and Technology
  • Weifang University
  • Shanghai Artificial Intelligence Laboratory
  • Brunel University London
  • University of Barcelona

科研成果: 期刊稿件文章同行评审

3 引用 (Scopus)

摘要

Remote Sensing Image-Text Retrieval (RSITR) aims to bridge the semantic gap between heterogeneous modalities and plays a vital role in various geospatial applications. As a lower dimensional and more concise modality than image, the text modality is more discriminative and may dominate the optimization process. The nonnegligible imbalanced cross-modal optimization remains a bottleneck to enhancing the model's performance. To address this issue, this study proposes a Representation Discrepancy Bridging (RDB) method for the RSITR task. On the one hand, a Cross-Modal Asymmetric Adapter (CMAA) is designed to enable modality-specific optimization and improve feature alignment. The CMAA comprises a Visual Enhancement Adapter (VEA) and a Text Semantic Adapter (TSA). VEA mines fine-grained image features through Differential Attention (DA) mechanism, while TSA identifies key textual semantics through Hierarchical Attention (HA) mechanism. On the other hand, this study extends the traditional single-task retrieval framework to a dual-task optimization framework and develops a Dual-Task Consistency Loss (DTCL). The DTCL improves cross-modal alignment robustness through an adaptive weighted combination of cross-modal, classification, and exponential moving average consistency constraints. Experiments on RSICD and RSITMD datasets show that the proposed RDB method achieves a 6 %–11 % improvement in mR metrics compared to state-of-the-art Parameter-Efficient Fine-Tuning (PEFT) methods and a 1.15 %–2 % improvement over the fully fine-tuned GeoRSCLIP model.

源语言英语
文章编号130915
期刊Neurocomputing
650
DOI
出版状态已出版 - 14 10月 2025
已对外发布

指纹

探究 'Representation discrepancy bridging method for remote sensing image-text retrieval' 的科研主题。它们共同构成独一无二的指纹。

引用此