Abstract
Multimodal remote sensing semantic segmentation based on optical and digital surface model (Opt-DSM) data is pivotal for comprehensive scene interpretation. However, prevailing methodologies often lack a unified vision foundation model and encounter significant challenges in bridging modality gaps and achieving effective feature fusion. Conventional models, such as the segment anything model (SAM), exhibit inherent limitations when addressing the unique complexities of multimodal remote sensing, particularly in managing cross-modal discrepancies and intricate surface structures. In this study, we present vision foundation model-driven multiscale expert tuning (VF-MET), an innovative framework meticulously tailored for Opt-DSM semantic segmentation tasks. VF-MET incorporates an adaptive multiscale expert tuning (AMET) strategy, which substantially enhances the feature extraction capabilities of vision foundation models. This enables the robust capture of cross-scale and morphologically irregular objects, while simultaneously preserving superior generalization ability. To further address the segmentation of densely distributed and weakly correlated regions, we propose a collaborative box-point prompt mechanism (CBPM), which significantly improves spatial localization and contextual discrimination. Moreover, we introduce a two-stage mask decoder (TSMD) that facilitates efficient multimodal feature fusion and augments contextual understanding. Extensive experiments conducted on public Opt-DSM benchmark datasets unequivocally demonstrate that VF-MET achieves state-of-the-art performance. Comprehensive ablation studies further substantiate the indispensable contributions of each constituent module within the proposed architecture.
| Original language | English |
|---|---|
| Article number | 5652817 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 63 |
| DOIs | |
| State | Published - 2025 |
Keywords
- Feature fusion
- frequency-aware modeling
- remote sensing imagery
- road extraction
- state-space model
Fingerprint
Dive into the research topics of 'Vision Foundation Model-Driven Multiscale Expert Tuning for Multimodal Remote Sensing Semantic Segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver