TY - JOUR
T1 - Interactive Fusion of Embodied Multimodal Data for Smart Cockpit Perception
AU - Fang, Aiqing
AU - Jiang, Yutong
AU - Lu, Nan
AU - Li, Ying
N1 - Publisher Copyright:
© 2026, China Ordnance Industry Corporation. All rights reserved.
PY - 2026/6
Y1 - 2026/6
N2 - Multimodal data fusion, as a key supporting technology for embodied intelligence, plays an important role in embodied interaction tasks such as autonomous driving, aircraft visual navigation, and military platforms. However, the traditional fusion paradigms face problems such as the difficulty in ensuring the fusion quality, the insufficient model generalization, and the lack of interpretability in decision-making processes in complex interactive environments. Therefore, this paper proposes a visual language-guided embodied multimodal data interactive fusion method for the multispectral environment perception and driver gaze interaction scenarios of intelligent agent vehicles in the military field. On the one hand, an interpretable multi-scale interactive fusion model is designed based on the embodied cognition theory and the Kolmogorov Arnold representation theorem. It is used to jointly model the multispectral environmental data and driver eye movement data, achieving the fusion enhancement of the focused area. On this basis, the semantic prior and representation ability of the visual language pretrained large model are introduced to optimize the multimodal fusion small model. The experimental results show that the proposed method overcomes the problems of difficult generalization and insufficient interpretability of multimodal fusion in complex interactive environments to a certain extent, and significantly improves the clarity of fused images, providing technical support for the application of interpretable human-machine interaction and embodied intelligence in transparent smart cockpits.
AB - Multimodal data fusion, as a key supporting technology for embodied intelligence, plays an important role in embodied interaction tasks such as autonomous driving, aircraft visual navigation, and military platforms. However, the traditional fusion paradigms face problems such as the difficulty in ensuring the fusion quality, the insufficient model generalization, and the lack of interpretability in decision-making processes in complex interactive environments. Therefore, this paper proposes a visual language-guided embodied multimodal data interactive fusion method for the multispectral environment perception and driver gaze interaction scenarios of intelligent agent vehicles in the military field. On the one hand, an interpretable multi-scale interactive fusion model is designed based on the embodied cognition theory and the Kolmogorov Arnold representation theorem. It is used to jointly model the multispectral environmental data and driver eye movement data, achieving the fusion enhancement of the focused area. On this basis, the semantic prior and representation ability of the visual language pretrained large model are introduced to optimize the multimodal fusion small model. The experimental results show that the proposed method overcomes the problems of difficult generalization and insufficient interpretability of multimodal fusion in complex interactive environments to a certain extent, and significantly improves the clarity of fused images, providing technical support for the application of interpretable human-machine interaction and embodied intelligence in transparent smart cockpits.
KW - Kolmogorov-Arnold network
KW - data fusion
KW - embodied intelligence
KW - frequency domain transformation
UR - https://www.scopus.com/pages/publications/105043840562
U2 - 10.12382/bgxb.2025.0987
DO - 10.12382/bgxb.2025.0987
M3 - 文章
AN - SCOPUS:105043840562
SN - 1000-1093
VL - 47
JO - Binggong Xuebao/Acta Armamentarii
JF - Binggong Xuebao/Acta Armamentarii
IS - 6
ER -