TY - JOUR
T1 - A Multimodal Deep Learning Framework for Spatial Room Impulse Response Generation in VR Auralization
AU - Li, Zhiyu
AU - Zhao, Xinpei
AU - Wang, Jing
AU - Li, Junfeng
AU - Chen, Jingdong
N1 - Publisher Copyright:
© 2025 IEEE. All rights reserved,
PY - 2026
Y1 - 2026
N2 - Online auralization in virtual reality (VR) requires fast yet perceptually accurate synthesis of spatial room impulse responses (SRIRs) under complex geometries, frequency-dependent materials, and dynamically changing source-listener configurations. Existing learning-based methods generally concentrate on generating only the short early-reflection portion of SRIRs, while full-response simulation with geometric acoustics (GA) incurs substantial computational cost at high reflection orders, leading to a trade-off between acoustic fidelity and practical runtime efficiency. In this work, we propose SRERS, a scene-waveform multimodal framework that generates full-length SRIRs from a face-based scene representation with acoustic attributes, source-listener coordinates, and a low-order reflection (LoR) computed via GA. The LoR is incorporated as an auxiliary modality to provide a physically grounded temporal anchor for sparse early arrivals, enabling the network to learn residual components beyond the LoR that model scattering, occlusion, and frequency-dependent coloration. To support training and generalization under diverse acoustic conditions, we further construct a dedicated SRIR dataset with enhanced variability. Experimental results demonstrate that SRERS consistently outperforms state-of-the-art baselines in both full-length and 4096-sample early-reflection SRIR generation. The performance gains remain evident under a strict w/o LoR protocol that excludes deterministic LoR contributions. Complexity analysis shows that SRERS enables online SRIR updating with controllable computational cost, while subjective listening evaluations validate improved perceptual similarity and robustness across conditions.
AB - Online auralization in virtual reality (VR) requires fast yet perceptually accurate synthesis of spatial room impulse responses (SRIRs) under complex geometries, frequency-dependent materials, and dynamically changing source-listener configurations. Existing learning-based methods generally concentrate on generating only the short early-reflection portion of SRIRs, while full-response simulation with geometric acoustics (GA) incurs substantial computational cost at high reflection orders, leading to a trade-off between acoustic fidelity and practical runtime efficiency. In this work, we propose SRERS, a scene-waveform multimodal framework that generates full-length SRIRs from a face-based scene representation with acoustic attributes, source-listener coordinates, and a low-order reflection (LoR) computed via GA. The LoR is incorporated as an auxiliary modality to provide a physically grounded temporal anchor for sparse early arrivals, enabling the network to learn residual components beyond the LoR that model scattering, occlusion, and frequency-dependent coloration. To support training and generalization under diverse acoustic conditions, we further construct a dedicated SRIR dataset with enhanced variability. Experimental results demonstrate that SRERS consistently outperforms state-of-the-art baselines in both full-length and 4096-sample early-reflection SRIR generation. The performance gains remain evident under a strict w/o LoR protocol that excludes deterministic LoR contributions. Complexity analysis shows that SRERS enables online SRIR updating with controllable computational cost, while subjective listening evaluations validate improved perceptual similarity and robustness across conditions.
KW - Spatial room impulse response
KW - auralization
KW - multimodal deep learning
KW - virtual reality
UR - https://www.scopus.com/pages/publications/105046340816
U2 - 10.1109/TASLPRO.2026.3717240
DO - 10.1109/TASLPRO.2026.3717240
M3 - 文章
AN - SCOPUS:105046340816
SN - 2998-4173
VL - 34
SP - 3854
EP - 3869
JO - IEEE Transactions on Audio, Speech and Language Processing
JF - IEEE Transactions on Audio, Speech and Language Processing
ER -