Skip to main navigation Skip to search Skip to main content

Weakly Paired Remote Sensing Change Captioning With Large Multimodal Model

  • Northwestern Polytechnical University Xian
  • Northwest Institute of Nuclear Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Automatically generating textual descriptions of changes between bi-temporal remote sensing images remains challenging due to the labor-intensive and time-consuming process of obtaining fully aligned image-caption paired data. To address this limitation, we propose a novel weakly paired learning framework that harnesses the powerful visual and language capabilities of large multimodal models (LMMs) for remote sensing change captioning. Unlike conventional methods that rely on precisely annotated image-caption pairs, our approach effectively utilizes weakly paired tags that encapsulate changed objects and their corresponding actions for each image. The proposed framework is built upon an LMM architecture and operates synergistically across two distinct visual-language phases via parameter-efficient adaptation. In the visual change perception phase, we characterize spatiotemporal variations by identifying cross-temporal object changes and their associated actions through tag recognition. Subsequently, in the language interpretation phase, we leverage the pretrained LMM to convert these visual change tags into coherent and semantically rich textual descriptions, exploiting the linguistic knowledge embedded in the LMM. To further enhance the alignment of LMM-generated captions with human annotations, we introduce a train-time expansion strategy, which refines alignment by rewriting synonymously expressed captions back into their original phrasing. This strategy improves alignment without incurring additional inference costs. As a plug-and-play method, our framework is compatible with various LMMs and achieves performance comparable to state-of-the-art fully supervised approaches, demonstrating its effectiveness in remote sensing change captioning. The code will be available at https://github.com/mrazhou/WeCap.

Original languageEnglish
Article number4410310
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume64
DOIs
StatePublished - 2026

Keywords

  • Change captioning
  • large multimodal model (LMM)
  • remote sensing
  • weakly paired

Fingerprint

Dive into the research topics of 'Weakly Paired Remote Sensing Change Captioning With Large Multimodal Model'. Together they form a unique fingerprint.

Cite this