Abstract
Remote sensing change detection (RSCD) plays a crucial role in applications such as environmental monitoring and urban planning. With the emergence of foundational vision-language models like CLIP, there is growing interest in integrating textual information into vision tasks. However, in the RSCD domain, limited efforts have been made to effectively leverage textual cues, and challenges persist in capturing differential features. To address these issues, this study proposes the multimodal difference augmentation learning change detection network (MdaCD) that fully exploits textual information and enhances differential feature learning (DFL). MdaCD introduces a CLIP-guided masking process to direct textual descriptions toward image differences, and a multimodal fusion and difference augmentation process to integrate and refine differential features across modalities. The CLIP-guided masking process applies masking to bitemporal image pairs before generating text prompts, enabling a more targeted analysis of changes. Meanwhile, the multimodal fusion and difference augmentation process computes a fused attention map to integrate visual and textual cues, effectively amplifying relevant differences. By applying difference augmentation functions, the differential features from both visual and textual embeddings are further refined and strengthened. The effectiveness of the proposed MdaCD model is validated through extensive experiments on two public RSCD datasets, where it achieves state-of-the-art performance with IoU scores of 84.88% on LEVIR-CD and 71.96% on SYSU-CD.
| Original language | English |
|---|---|
| Article number | 5652111 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 63 |
| DOIs | |
| State | Published - 2025 |
Keywords
- Change detection (CD)
- differential feature learning (DFL)
- multimodal learning
- remote sensing
Fingerprint
Dive into the research topics of 'Multimodal Difference Augmentation Learning for Remote Sensing Change Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver