Skip to main navigation Skip to search Skip to main content

Vision Foundation Model-Driven Multiscale Expert Tuning for Multimodal Remote Sensing Semantic Segmentation

  • Jiayuan Li
  • , Zhen Wang
  • , Nan Xu
  • , Zhuhong You
  • , De Shuang Huang
  • Northwestern Polytechnical University Xian
  • Xijing University
  • Shenzhen University
  • Guangxi Academy of Agricultural Sciences

Research output: Contribution to journalArticlepeer-review

5 Scopus citations

Abstract

Multimodal remote sensing semantic segmentation based on optical and digital surface model (Opt-DSM) data is pivotal for comprehensive scene interpretation. However, prevailing methodologies often lack a unified vision foundation model and encounter significant challenges in bridging modality gaps and achieving effective feature fusion. Conventional models, such as the segment anything model (SAM), exhibit inherent limitations when addressing the unique complexities of multimodal remote sensing, particularly in managing cross-modal discrepancies and intricate surface structures. In this study, we present vision foundation model-driven multiscale expert tuning (VF-MET), an innovative framework meticulously tailored for Opt-DSM semantic segmentation tasks. VF-MET incorporates an adaptive multiscale expert tuning (AMET) strategy, which substantially enhances the feature extraction capabilities of vision foundation models. This enables the robust capture of cross-scale and morphologically irregular objects, while simultaneously preserving superior generalization ability. To further address the segmentation of densely distributed and weakly correlated regions, we propose a collaborative box-point prompt mechanism (CBPM), which significantly improves spatial localization and contextual discrimination. Moreover, we introduce a two-stage mask decoder (TSMD) that facilitates efficient multimodal feature fusion and augments contextual understanding. Extensive experiments conducted on public Opt-DSM benchmark datasets unequivocally demonstrate that VF-MET achieves state-of-the-art performance. Comprehensive ablation studies further substantiate the indispensable contributions of each constituent module within the proposed architecture.

Original languageEnglish
Article number5652817
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume63
DOIs
StatePublished - 2025

Keywords

  • Feature fusion
  • frequency-aware modeling
  • remote sensing imagery
  • road extraction
  • state-space model

Fingerprint

Dive into the research topics of 'Vision Foundation Model-Driven Multiscale Expert Tuning for Multimodal Remote Sensing Semantic Segmentation'. Together they form a unique fingerprint.

Cite this