Abstract
Remote sensing image change captioning aims to generate natural language descriptions of dynamic changes between bitemporal images. This task faces challenges of aligning high-level semantics with sparse change regions and modeling long-range dependencies in diverse change patterns. To alleviate these limitations, we propose the spatial-semantic alignment and change-aware network, dubbed SACNet. It introduces a semantic-spatial collaborative encoding module synergizing a CNN-based spatial encoder with a CLIP-based semantic encoder, using precise spatial cues to refine global semantic features. A change-driven adaptive feature calibration module emphasizes discriminative change signals via intertemporal feature dissimilarity modeling, suppressing background noise and highlighting key changes. In addition, a multipath state space model (SSM)-guided spatial representation module is designed to capture diverse complex spatial change topologies. Extensive experiments on the LEVIR-CC and WHU-CDC datasets demonstrate that SACNet achieves state-of-the-art performance, generating both high-precision textual descriptions and approximate change masks. Code will be available at https://github.com/CVer-Yang/SACNet.
| Original language | English |
|---|---|
| Article number | 5621611 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 64 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Change captioning
- key change features
- multitask collaborative optimization
- remote sensing
- state space models (SSMs)
Fingerprint
Dive into the research topics of 'Spatial-Semantic Alignment and Change-Aware Network for Remote Sensing Image Change Captioning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver