跳到主要导航 跳到搜索 跳到主要内容

DIA: Deriving linguistic information from auxiliary languages for remote sensing image captioning

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

7 引用 (Scopus)

摘要

Remote sensing image captioning (RSIC) is a cross-modal task aimed at describing scene categories, object classes, and their spatial relationships in remote sensing images using natural language. Existing methods typically focus on training models in single language, neglecting the linguistic-enhancing information derived from syntactic structure differences and diverse expressions of the same objects and scenes. This information, present in cross-linguistic annotated data, can significantly enhance language perception and enrich training data. To verify the effectiveness of this information, we propose an auxiliary language-enhanced network called DIA, which leverages linguistic information from auxiliary languages to improve the quality and fluency of target language generation. DIA consists of shared visual feature extractor, target language generator, and auxiliary language generator. The shared visual feature extractor integrates the Linguistic-Irrelevant Feature Enrichment (LiFE) module, while a Linguistic Bridge connects the target and auxiliary language generators. The LiFE module employs linguistic-irrelevant feature extraction and multi-view attention to extract precise visual features, enriching the representations while minimizing language bias. Multi-view attention balances deep semantic expressions and linguistic-irrelevant features. The Linguistic Bridge establishes interactive pathway between the target language generator (ALG) and the auxiliary language generator (TLG), enabling the TLG to learn from the ALG's language modeling capabilities. This interaction allows the TLG to handle complex visual features, improving language generation performance. Extensive experiments demonstrate that our model achieves significant performance improvements on the UCM, Sydney, RSICD, and NWPU datasets. Specifically, on the UCM dataset, BLEU-4 is improved by 5.06 %, and CIDEr is improved by 16.86 %.

源语言英语
文章编号112209
期刊Pattern Recognition
171
DOI
出版状态已出版 - 3月 2026

学术指纹

探究 'DIA: Deriving linguistic information from auxiliary languages for remote sensing image captioning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此