TY - JOUR
T1 - Multi-scale fusion algorithm for infrared and visible images guided by semantic segmentation
AU - Yang, Jianhua
AU - Ruan, Shucheng
AU - Huo, Changling
AU - Guo, Quanming
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/8
Y1 - 2026/8
N2 - To address the issues of target feature attenuation and background interference in infrared–visible image fusion under complex scenarios, this paper proposes a multi-scale adaptive fusion algorithm guided by semantic segmentation. First, visible images undergo enhancement preprocessing through median filtering, contrast-limited adaptive histogram equalization (CLAHE), and unsharp masking. Guided filters then decompose infrared and visible images into base and detail layers. Simultaneously, an enhanced collaborative attention-based lightweight segmentation network (SA-DeepLab) performs semantic segmentation on the infrared image, generating high-confidence object region masks to provide precise guidance for subsequent fusion. During fusion, a hierarchical adaptive strategy is proposed: base layer fusion preserves infrared structures in target regions while integrating visible light information into background areas based on semantic masks; detail layer fusion constructs multi-feature adaptive weights combining saliency, gradient, and Laplacian energy to intelligently fuse texture details. Finally, post-processing with color preservation and contrast enhancement outputs a fused image that highlights thermal targets while maintaining natural colors and rich details. Experiments on the MSRS, RoadScene and TNO datasets demonstrate that the proposed algorithm outperforms existing mainstream methods in both subjective visual quality and objective evaluation metrics. It exhibits strong robustness and practicality in complex driving scenarios such as nighttime, low-light conditions, and occlusions, providing reliable environmental perception support for advanced driver assistance systems.
AB - To address the issues of target feature attenuation and background interference in infrared–visible image fusion under complex scenarios, this paper proposes a multi-scale adaptive fusion algorithm guided by semantic segmentation. First, visible images undergo enhancement preprocessing through median filtering, contrast-limited adaptive histogram equalization (CLAHE), and unsharp masking. Guided filters then decompose infrared and visible images into base and detail layers. Simultaneously, an enhanced collaborative attention-based lightweight segmentation network (SA-DeepLab) performs semantic segmentation on the infrared image, generating high-confidence object region masks to provide precise guidance for subsequent fusion. During fusion, a hierarchical adaptive strategy is proposed: base layer fusion preserves infrared structures in target regions while integrating visible light information into background areas based on semantic masks; detail layer fusion constructs multi-feature adaptive weights combining saliency, gradient, and Laplacian energy to intelligently fuse texture details. Finally, post-processing with color preservation and contrast enhancement outputs a fused image that highlights thermal targets while maintaining natural colors and rich details. Experiments on the MSRS, RoadScene and TNO datasets demonstrate that the proposed algorithm outperforms existing mainstream methods in both subjective visual quality and objective evaluation metrics. It exhibits strong robustness and practicality in complex driving scenarios such as nighttime, low-light conditions, and occlusions, providing reliable environmental perception support for advanced driver assistance systems.
KW - Attention mechanism
KW - Infrared and visible image fusion
KW - Intelligent perception
KW - Semantic segmentation
UR - https://www.scopus.com/pages/publications/105039288079
U2 - 10.1016/j.infrared.2026.106648
DO - 10.1016/j.infrared.2026.106648
M3 - 文章
AN - SCOPUS:105039288079
SN - 1350-4495
VL - 157
JO - Infrared Physics and Technology
JF - Infrared Physics and Technology
M1 - 106648
ER -