TY - GEN
T1 - A Multi-Scale Multimodal Fusion Method with Attention Mechanism for Constrained Environments
AU - Wang, Hao
AU - Xia, Tian
AU - Wang, Zehan
AU - Jiang, Pengfei
AU - Yang, Beiya
AU - Shi, Haobin
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Multi-Temperature environments in industrial intelligent warehouses facilitate the formation of salt-fog, which degrades the performance of both vision and LiDAR sensors. Thus, the salt-fog severely undermines the reliability of key tasks such as object detection and trajectory prediction. To address this challenge, we propose a Multi-Scale Multi-Modal Fusion Method with Attention Mechanism (MMA). The proposed approach consists of two core modules. First, an attention-based fusion module performs interaction between visual and LiDAR features in a unified feature space, realizing information complementarity across modalities. Second, a multi-scale feature enhancement module reconstructs and strengthens the fused features, alleviating local feature loss and semantic degradation caused by salt-fog interference. In addition, an intelligent warehouse simulation environment with controllable salt-fog intensity and dynamic entities is constructed, and the proposed method is systematically evaluated on two representative downstream tasks: object detection and trajectory prediction. Experimental results demonstrate that MMA significantly improves perception performance and prediction stability in varying downstream tasks. Under severe salt-fog conditions, mAP@50:95 for object detection is improved by 2.9%, and the average displacement error for trajectory prediction is reduced by 30% compared to the baseline.
AB - Multi-Temperature environments in industrial intelligent warehouses facilitate the formation of salt-fog, which degrades the performance of both vision and LiDAR sensors. Thus, the salt-fog severely undermines the reliability of key tasks such as object detection and trajectory prediction. To address this challenge, we propose a Multi-Scale Multi-Modal Fusion Method with Attention Mechanism (MMA). The proposed approach consists of two core modules. First, an attention-based fusion module performs interaction between visual and LiDAR features in a unified feature space, realizing information complementarity across modalities. Second, a multi-scale feature enhancement module reconstructs and strengthens the fused features, alleviating local feature loss and semantic degradation caused by salt-fog interference. In addition, an intelligent warehouse simulation environment with controllable salt-fog intensity and dynamic entities is constructed, and the proposed method is systematically evaluated on two representative downstream tasks: object detection and trajectory prediction. Experimental results demonstrate that MMA significantly improves perception performance and prediction stability in varying downstream tasks. Under severe salt-fog conditions, mAP@50:95 for object detection is improved by 2.9%, and the average displacement error for trajectory prediction is reduced by 30% compared to the baseline.
KW - Attention
KW - Constrained Environment
KW - Industrial Warehouse
KW - Multi-Scale Feature
KW - Multimodal Fusion
UR - https://www.scopus.com/pages/publications/105044787295
U2 - 10.1109/RAITS68656.2026.11580254
DO - 10.1109/RAITS68656.2026.11580254
M3 - 会议稿件
AN - SCOPUS:105044787295
T3 - Proceedings - 2026 International Conference on Robotics, Automation and Intelligent Transportation Systems, RAITS 2026
BT - Proceedings - 2026 International Conference on Robotics, Automation and Intelligent Transportation Systems, RAITS 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 International Conference on Robotics, Automation and Intelligent Transportation Systems, RAITS 2026
Y2 - 23 January 2026 through 25 January 2026
ER -