TY - JOUR
T1 - AutoRoadSAM
T2 - Multimodal Remote Sensing Road Extraction with Structure-Semantic Awareness via Auto-Prompting Vision Foundation Models
AU - Li, Jiayuan
AU - Wang, Zhen
AU - Sun, Xiao
AU - Lv, Zhiyong
AU - Xu, Nan
AU - You, Zhuhong
AU - Huang, De Shuang
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - The integration of multimodal data holds great promise for advancing road extraction in remote sensing. However, existing approaches are limited by the lack of unified end-to-end frameworks for diverse modality combinations, suboptimal multimodal feature fusion, and challenges in capturing the slender, winding, and complex topological structures of roads. In this article, we propose AutoRoadSAM, a novel end-to-end framework for multimodal road extraction that fully exploits the powerful visual representation capabilities of the segment anything model (SAM) and, for the first time, introduces an auto-prompting mechanism via a dynamic snake convolution-based decoder. This decoder adaptively generates task-specific prompts by capturing fine-grained local geometric features from auxiliary modality branches, enabling precise alignment with complex road structures. To further enhance multimodal feature fusion and topological perception, we design the cross-modal information interaction (CMII) module, which facilitates global context modeling and cross-modal interaction, while strengthening the representation of intricate road topology through multidirectional snake scanning. Moreover, we incorporate a mask decoder with cross-polarity-aware linear attention (CPLAM) to boost decoding efficiency and effectively address pixel imbalance. Together, these innovations enable AutoRoadSAM to achieve superior structure- and semantic-aware road extraction across diverse modality combinations. Extensive experiments on six public datasets and four modality combinations demonstrate that AutoRoadSAM consistently outperforms state-of-the-art methods, validating the effectiveness and generalization capability of each proposed component. The code is available at https://github.com/NWPUFranklee/AutoRoadSAM.git.
AB - The integration of multimodal data holds great promise for advancing road extraction in remote sensing. However, existing approaches are limited by the lack of unified end-to-end frameworks for diverse modality combinations, suboptimal multimodal feature fusion, and challenges in capturing the slender, winding, and complex topological structures of roads. In this article, we propose AutoRoadSAM, a novel end-to-end framework for multimodal road extraction that fully exploits the powerful visual representation capabilities of the segment anything model (SAM) and, for the first time, introduces an auto-prompting mechanism via a dynamic snake convolution-based decoder. This decoder adaptively generates task-specific prompts by capturing fine-grained local geometric features from auxiliary modality branches, enabling precise alignment with complex road structures. To further enhance multimodal feature fusion and topological perception, we design the cross-modal information interaction (CMII) module, which facilitates global context modeling and cross-modal interaction, while strengthening the representation of intricate road topology through multidirectional snake scanning. Moreover, we incorporate a mask decoder with cross-polarity-aware linear attention (CPLAM) to boost decoding efficiency and effectively address pixel imbalance. Together, these innovations enable AutoRoadSAM to achieve superior structure- and semantic-aware road extraction across diverse modality combinations. Extensive experiments on six public datasets and four modality combinations demonstrate that AutoRoadSAM consistently outperforms state-of-the-art methods, validating the effectiveness and generalization capability of each proposed component. The code is available at https://github.com/NWPUFranklee/AutoRoadSAM.git.
KW - Auto-prompting
KW - feature fusion
KW - multimodal remote sensing
KW - road extraction
KW - vision foundation models
UR - https://www.scopus.com/pages/publications/105028902538
U2 - 10.1109/TGRS.2026.3658664
DO - 10.1109/TGRS.2026.3658664
M3 - 文章
AN - SCOPUS:105028902538
SN - 0196-2892
VL - 64
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 5607617
ER -