TY - JOUR
T1 - Multiscale Dynamic Attention and Transformer-Driven Interactive Multifeature Fusion Approach for Advanced Crack Detection and Segmentation
AU - Zaheer, Qasim
AU - Atta, Zunaira
AU - Ehsan, Haleema
AU - Malik, Momina
AU - Shah, Syed Rehan
AU - Long, Xu
N1 - Publisher Copyright:
© 2026 American Society of Civil Engineers.
PY - 2026/9/1
Y1 - 2026/9/1
N2 - Salient Object Detection is a key preprocessing technique for identifying prominent regions in images, but its use in pavement surface crack detection is limited due to scale and feature differences. To address this, we introduce the multiscale dynamic attention and crack detection (MDACD) model as an innovative solution for crack detection and segmentation. MDACD integrates a Swin transformer backbone, decoder blocks, and dynamic fusion layers (DFLs) to capture global contextual features, refine spatial details, and improve the network's ability to handle irregular crack patterns. By optimizing fusion layers for enhanced feature interaction, MDACD achieves a robust balance between accuracy (97.03%), weighted average recall (97%), weighted average precision (98%), and weighted average F1-score (97%). The model incorporates multiple loss functions alongside dice and intersection over union (IoU) metrics for comprehensive performance evaluation. Comparative analysis on a publicly available benchmark data set confirms that MDACD surpasses state-of-the-art models, demonstrating superior effectiveness in crack detection.
AB - Salient Object Detection is a key preprocessing technique for identifying prominent regions in images, but its use in pavement surface crack detection is limited due to scale and feature differences. To address this, we introduce the multiscale dynamic attention and crack detection (MDACD) model as an innovative solution for crack detection and segmentation. MDACD integrates a Swin transformer backbone, decoder blocks, and dynamic fusion layers (DFLs) to capture global contextual features, refine spatial details, and improve the network's ability to handle irregular crack patterns. By optimizing fusion layers for enhanced feature interaction, MDACD achieves a robust balance between accuracy (97.03%), weighted average recall (97%), weighted average precision (98%), and weighted average F1-score (97%). The model incorporates multiple loss functions alongside dice and intersection over union (IoU) metrics for comprehensive performance evaluation. Comparative analysis on a publicly available benchmark data set confirms that MDACD surpasses state-of-the-art models, demonstrating superior effectiveness in crack detection.
KW - Crack detection
KW - Crack segmentation
KW - Pixel level segmentation
KW - Salient object detection
UR - https://www.scopus.com/pages/publications/105041038253
U2 - 10.1061/JITSE4.ISENG-2912
DO - 10.1061/JITSE4.ISENG-2912
M3 - 文章
AN - SCOPUS:105041038253
SN - 1076-0342
VL - 32
JO - Journal of Infrastructure Systems
JF - Journal of Infrastructure Systems
IS - 3
M1 - 04026012
ER -