Abstract
Accurate semantic segmentation of remote sensing images is crucial for urban planning, land monitoring, and related applications, yet remains challenging due to the heterogeneity of multimodal data and limited target attribute representation. In this work, we propose a novel Multimodal Diffusion Prior guided Self-Text Attention Network (MDPNet) that introduces three key innovations: a denoising diffusion probabilistic model to extract robust structural priors from digital surface model data, effectively suppressing noise and enhancing fine-grained boundary features; a multimodal prior feature guidance module that employs a cross-modal selective state-space mechanism and a dislocated stacking strategy, enabling explicit patch-level fusion and capturing long-range dependencies between modalities; and a self-text attention mechanism that automatically generates and aligns category-related textual cues from segmentation masks, eliminating the need for external textual input and providing fine-grained semantic guidance. Extensive experiments on the ISPRS Potsdam and Vaihingen public datasets demonstrate that MDPNet sets a new state-of-the-art, particularly excelling in building edge delineation, robustness under complex scenarios, and balanced segmentation across multiple categories. Our approach outperforms existing mainstream methods in both overall accuracy and resilience, and provides a text-free, end-to-end solution for multimodal remote sensing semantic segmentation.
| Original language | English |
|---|---|
| Pages (from-to) | 3924-3942 |
| Number of pages | 19 |
| Journal | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| Volume | 19 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Remote sensing image (RSI)
- diffusion probabilistic model
- multimodal fusion
- self-text attention
- semantic segmentation
- state-space model
Fingerprint
Dive into the research topics of 'MDPNet: Multimodal Diffusion Prior Guided Self-Text Attention Network for Remote Sensing Semantic Segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver