TY - JOUR
T1 - Weakly Paired Remote Sensing Change Captioning With Large Multimodal Model
AU - Zhou, Qing
AU - Wu, Junzheng
AU - Jia, Yuyu
AU - Zhou, Shihao
AU - Gao, Junyu
AU - Ni, Weiping
AU - Wang, Qi
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Automatically generating textual descriptions of changes between bi-temporal remote sensing images remains challenging due to the labor-intensive and time-consuming process of obtaining fully aligned image-caption paired data. To address this limitation, we propose a novel weakly paired learning framework that harnesses the powerful visual and language capabilities of large multimodal models (LMMs) for remote sensing change captioning. Unlike conventional methods that rely on precisely annotated image-caption pairs, our approach effectively utilizes weakly paired tags that encapsulate changed objects and their corresponding actions for each image. The proposed framework is built upon an LMM architecture and operates synergistically across two distinct visual-language phases via parameter-efficient adaptation. In the visual change perception phase, we characterize spatiotemporal variations by identifying cross-temporal object changes and their associated actions through tag recognition. Subsequently, in the language interpretation phase, we leverage the pretrained LMM to convert these visual change tags into coherent and semantically rich textual descriptions, exploiting the linguistic knowledge embedded in the LMM. To further enhance the alignment of LMM-generated captions with human annotations, we introduce a train-time expansion strategy, which refines alignment by rewriting synonymously expressed captions back into their original phrasing. This strategy improves alignment without incurring additional inference costs. As a plug-and-play method, our framework is compatible with various LMMs and achieves performance comparable to state-of-the-art fully supervised approaches, demonstrating its effectiveness in remote sensing change captioning. The code will be available at https://github.com/mrazhou/WeCap.
AB - Automatically generating textual descriptions of changes between bi-temporal remote sensing images remains challenging due to the labor-intensive and time-consuming process of obtaining fully aligned image-caption paired data. To address this limitation, we propose a novel weakly paired learning framework that harnesses the powerful visual and language capabilities of large multimodal models (LMMs) for remote sensing change captioning. Unlike conventional methods that rely on precisely annotated image-caption pairs, our approach effectively utilizes weakly paired tags that encapsulate changed objects and their corresponding actions for each image. The proposed framework is built upon an LMM architecture and operates synergistically across two distinct visual-language phases via parameter-efficient adaptation. In the visual change perception phase, we characterize spatiotemporal variations by identifying cross-temporal object changes and their associated actions through tag recognition. Subsequently, in the language interpretation phase, we leverage the pretrained LMM to convert these visual change tags into coherent and semantically rich textual descriptions, exploiting the linguistic knowledge embedded in the LMM. To further enhance the alignment of LMM-generated captions with human annotations, we introduce a train-time expansion strategy, which refines alignment by rewriting synonymously expressed captions back into their original phrasing. This strategy improves alignment without incurring additional inference costs. As a plug-and-play method, our framework is compatible with various LMMs and achieves performance comparable to state-of-the-art fully supervised approaches, demonstrating its effectiveness in remote sensing change captioning. The code will be available at https://github.com/mrazhou/WeCap.
KW - Change captioning
KW - large multimodal model (LMM)
KW - remote sensing
KW - weakly paired
UR - https://www.scopus.com/pages/publications/105040950203
U2 - 10.1109/TGRS.2026.3698503
DO - 10.1109/TGRS.2026.3698503
M3 - 文章
AN - SCOPUS:105040950203
SN - 0196-2892
VL - 64
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 4410310
ER -