Skip to main navigation Skip to search Skip to main content

MSCANet: Multi-scale Cross-Attention Fusion Network for Multimodal Remote Sensing Image Semantic Segmentation

  • Xiao Sun
  • , Jiayuan Li
  • , Zhen Wang
  • , Zhu Hong You
  • , Yu Li
  • , Zhe Zhang
  • Northwestern Polytechnical University Xian
  • Tianjin University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Semantic segmentation of remote sensing images is critical for land cover classification and disaster monitoring, but single-modal sources like optical or SAR imagery suffer from inherent limitations. To address this, we propose MSCANet, a multi-scale cross-attention fusion network for multimodal segmentation. MSCANet uses a dual-branch architecture, where optical and SAR images are processed by lightweight InceptionNext-Tiny encoders. At each stage, a Residual-Calibrated Three-Stage Attention (RCTA) module sequentially applies multi-head cross-attention to model inter-modal dependencies, followed by channel and spatial attention to recalibrate features and suppress modality-specific noise within a residual learning framework. This design ensures semantic alignment and robust information flow with linear complexity. Fused multi-scale features are then aggregated in a Multi-Scale Dilated-Skip Fusion (MSDSF) decoder, which uses serial dilated convolutions and dynamic skip connections to capture rich context and recover fine boundaries. Experiments on WHU-OPT-SAR and UBCv2 datasets show that MSCANet achieves state-of-the-art performance with low computational cost. Specifically, on WHU-OPT-SAR, MSCANet reaches 59.19% mIoU, outperforming the second-best method (CEN) by 2.35%, while requiring only 15.94 GFLOPs. These results demonstrate the effectiveness of MSCANet for robust and efficient multimodal remote sensing segmentation.

Original languageEnglish
Title of host publicationAdvanced Intelligent Computing Technology and Applications - 22nd International Conference on Intelligent Computing, ICIC 2026, Proceedings
EditorsDe-Shuang Huang, Qinhu Zhang, Yijie Pan, Chuanlei Zhang, Wei Chen, Bo Li, Wenzheng Bao, Prashan Premaratne
PublisherSpringer Science and Business Media Deutschland GmbH
Pages220-232
Number of pages13
ISBN (Print)9789819234028
DOIs
StatePublished - 2027
Event22nd International Conference on Intelligent Computing, ICIC 2026 - Toronto, Canada
Duration: 22 Jul 202626 Jul 2026

Publication series

NameLecture Notes in Computer Science
Volume16649 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference22nd International Conference on Intelligent Computing, ICIC 2026
Country/TerritoryCanada
CityToronto
Period22/07/2626/07/26

Keywords

  • Cross-Attention
  • Feature Representation
  • Lightweight Network
  • Multimodal Fusion
  • Remote Sensing Image
  • Semantic Segmentation

Fingerprint

Dive into the research topics of 'MSCANet: Multi-scale Cross-Attention Fusion Network for Multimodal Remote Sensing Image Semantic Segmentation'. Together they form a unique fingerprint.

Cite this