Abstract
The segment anything model (SAM) offers strong feature extraction and segmentation performance for remote sensing. However, for multimodal remote sensing semantic segmentation (MRSSS), SAM suffers from limited adaptation to domain-specific data, insufficient multimodal fusion, and ineffective cross-scale information interaction. To address these issues, we propose a unified end-to-end foundation framework with cross-scale calibration and adaptation fusion, termed multimodal remote sensing SAM (MRS-SAM), specifically designed for MRSSS tasks. MRS-SAM features three synergistic modules: first, a multimodal adaptive fine-tuning and feature enhancement module, which employs AdaLoRA-based parameter-efficient fine-tuning of ViT blocks for remote sensing adaptation, enables deep multimodal interaction through MulAdapter, and generates global multiscale pyramid features; second, an adjacent-scale multimodal fusion mechanism that enhances feature fusion via dual channel-spatial processing, effectively aligning and integrating heterogeneous modal information; and finally, a pyramid fusion Mamba module that leverages the efficient global sequence modeling of the state space model to facilitate cross-scale information interaction and eliminate semantic redundancy. Extensive experiments on three benchmark MRSSS datasets demonstrate that MRS-SAM consistently outperforms state-of-the-art methods across multiple quantitative evaluation metrics. Furthermore, ablation studies validate the effectiveness of each module in advancing multimodal feature fusion and calibration.
| Original language | English |
|---|---|
| Pages (from-to) | 19323-19339 |
| Number of pages | 17 |
| Journal | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| Volume | 19 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Adaptation fusion
- cross-scale calibration
- foundation model
- multimodal remote sensing (MRS)
- semantic segmentation
Fingerprint
Dive into the research topics of 'A Unified Cross-Scale Calibration and Adaptation Fusion Framework for Multimodal Remote Sensing Semantic Segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver