TY - JOUR
T1 - A Unified Cross-Scale Calibration and Adaptation Fusion Framework for Multimodal Remote Sensing Semantic Segmentation
AU - Deng, Shaoliang
AU - Li, Jiayuan
AU - Zhang, Yanyu
AU - Wang, Zhen
AU - You, Zhuhong
N1 - Publisher Copyright:
© 2008-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - The segment anything model (SAM) offers strong feature extraction and segmentation performance for remote sensing. However, for multimodal remote sensing semantic segmentation (MRSSS), SAM suffers from limited adaptation to domain-specific data, insufficient multimodal fusion, and ineffective cross-scale information interaction. To address these issues, we propose a unified end-to-end foundation framework with cross-scale calibration and adaptation fusion, termed multimodal remote sensing SAM (MRS-SAM), specifically designed for MRSSS tasks. MRS-SAM features three synergistic modules: first, a multimodal adaptive fine-tuning and feature enhancement module, which employs AdaLoRA-based parameter-efficient fine-tuning of ViT blocks for remote sensing adaptation, enables deep multimodal interaction through MulAdapter, and generates global multiscale pyramid features; second, an adjacent-scale multimodal fusion mechanism that enhances feature fusion via dual channel-spatial processing, effectively aligning and integrating heterogeneous modal information; and finally, a pyramid fusion Mamba module that leverages the efficient global sequence modeling of the state space model to facilitate cross-scale information interaction and eliminate semantic redundancy. Extensive experiments on three benchmark MRSSS datasets demonstrate that MRS-SAM consistently outperforms state-of-the-art methods across multiple quantitative evaluation metrics. Furthermore, ablation studies validate the effectiveness of each module in advancing multimodal feature fusion and calibration.
AB - The segment anything model (SAM) offers strong feature extraction and segmentation performance for remote sensing. However, for multimodal remote sensing semantic segmentation (MRSSS), SAM suffers from limited adaptation to domain-specific data, insufficient multimodal fusion, and ineffective cross-scale information interaction. To address these issues, we propose a unified end-to-end foundation framework with cross-scale calibration and adaptation fusion, termed multimodal remote sensing SAM (MRS-SAM), specifically designed for MRSSS tasks. MRS-SAM features three synergistic modules: first, a multimodal adaptive fine-tuning and feature enhancement module, which employs AdaLoRA-based parameter-efficient fine-tuning of ViT blocks for remote sensing adaptation, enables deep multimodal interaction through MulAdapter, and generates global multiscale pyramid features; second, an adjacent-scale multimodal fusion mechanism that enhances feature fusion via dual channel-spatial processing, effectively aligning and integrating heterogeneous modal information; and finally, a pyramid fusion Mamba module that leverages the efficient global sequence modeling of the state space model to facilitate cross-scale information interaction and eliminate semantic redundancy. Extensive experiments on three benchmark MRSSS datasets demonstrate that MRS-SAM consistently outperforms state-of-the-art methods across multiple quantitative evaluation metrics. Furthermore, ablation studies validate the effectiveness of each module in advancing multimodal feature fusion and calibration.
KW - Adaptation fusion
KW - cross-scale calibration
KW - foundation model
KW - multimodal remote sensing (MRS)
KW - semantic segmentation
UR - https://www.scopus.com/pages/publications/105040996016
U2 - 10.1109/JSTARS.2026.3697833
DO - 10.1109/JSTARS.2026.3697833
M3 - 文章
AN - SCOPUS:105040996016
SN - 1939-1404
VL - 19
SP - 19323
EP - 19339
JO - IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
JF - IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
ER -