TY - JOUR
T1 - DINO-MTP
T2 - Topology-Aware Multimodal Fusion of GPS Trajectories and Optical Imagery for Remote Sensing Road Extraction
AU - Yao, Yiming
AU - Chu, Peng
AU - Li, Jiayuan
AU - Wang, Zhen
AU - You, Zhuhong
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Accurate road extraction from remote sensing imagery is critical for urban mapping, navigation, and geospatial analysis, yet it remains highly challenging due to occlusions, thin and fragmented road structures, and the inherent limitations of single-modality data. To overcome these challenges, we propose DINO-MTP, a novel end-to-end framework that synergistically fuses GPS trajectory data with optical imagery for robust multimodal road extraction. Leveraging a self-supervised DINOv3 backbone, our approach introduces topology-aware and morphology-aware multimodal feature learning to explicitly address cross-modality heterogeneity, structural discontinuities, and background clutter. The framework incorporates a bidirectional attention alignment with modular attention masking to achieve deep semantic alignment and effective noise suppression within a unified embedding space. Additionally, it integrates spatial priors into channel attention to enhance the representation of elongated and interconnected road structures, while a feature-adaptive super-resolution pathway recovers fine details typically lost in conventional skip connections. An efficient dual-path cross-scale fusion module further ensures a coherent transition from global topology reconstruction to local geometric refinement. Extensive experiments and ablation studies on the BJRoad, RS-RGB-T and Porto datasets demonstrate that DINO-MTP consistently surpasses state-of-the-art unimodal and multimodal methods, delivering clearer, more continuous, and topologically faithful road network extraction in complex urban environments.
AB - Accurate road extraction from remote sensing imagery is critical for urban mapping, navigation, and geospatial analysis, yet it remains highly challenging due to occlusions, thin and fragmented road structures, and the inherent limitations of single-modality data. To overcome these challenges, we propose DINO-MTP, a novel end-to-end framework that synergistically fuses GPS trajectory data with optical imagery for robust multimodal road extraction. Leveraging a self-supervised DINOv3 backbone, our approach introduces topology-aware and morphology-aware multimodal feature learning to explicitly address cross-modality heterogeneity, structural discontinuities, and background clutter. The framework incorporates a bidirectional attention alignment with modular attention masking to achieve deep semantic alignment and effective noise suppression within a unified embedding space. Additionally, it integrates spatial priors into channel attention to enhance the representation of elongated and interconnected road structures, while a feature-adaptive super-resolution pathway recovers fine details typically lost in conventional skip connections. An efficient dual-path cross-scale fusion module further ensures a coherent transition from global topology reconstruction to local geometric refinement. Extensive experiments and ablation studies on the BJRoad, RS-RGB-T and Porto datasets demonstrate that DINO-MTP consistently surpasses state-of-the-art unimodal and multimodal methods, delivering clearer, more continuous, and topologically faithful road network extraction in complex urban environments.
KW - Multimodal fusion
KW - remote sensing
KW - road extraction
KW - self-supervised representation
KW - topology-aware
UR - https://www.scopus.com/pages/publications/105039566045
U2 - 10.1109/TGRS.2026.3695443
DO - 10.1109/TGRS.2026.3695443
M3 - 文章
AN - SCOPUS:105039566045
SN - 0196-2892
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
ER -