摘要
Remote sensing image change captioning aims to generate natural language descriptions of dynamic changes between bitemporal images. This task faces challenges of aligning high-level semantics with sparse change regions and modeling long-range dependencies in diverse change patterns. To alleviate these limitations, we propose the spatial-semantic alignment and change-aware network, dubbed SACNet. It introduces a semantic-spatial collaborative encoding module synergizing a CNN-based spatial encoder with a CLIP-based semantic encoder, using precise spatial cues to refine global semantic features. A change-driven adaptive feature calibration module emphasizes discriminative change signals via intertemporal feature dissimilarity modeling, suppressing background noise and highlighting key changes. In addition, a multipath state space model (SSM)-guided spatial representation module is designed to capture diverse complex spatial change topologies. Extensive experiments on the LEVIR-CC and WHU-CDC datasets demonstrate that SACNet achieves state-of-the-art performance, generating both high-precision textual descriptions and approximate change masks. Code will be available at https://github.com/CVer-Yang/SACNet.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 5621611 |
| 期刊 | IEEE Transactions on Geoscience and Remote Sensing |
| 卷 | 64 |
| DOI | |
| 出版状态 | 已出版 - 2026 |
指纹
探究 'Spatial-Semantic Alignment and Change-Aware Network for Remote Sensing Image Change Captioning' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver