跳到主要导航 跳到搜索 跳到主要内容

A Multimodal Deep Learning Framework for Spatial Room Impulse Response Generation in VR Auralization

  • Beijing Institute of Technology
  • CAS - Institute of Acoustics

科研成果: 期刊稿件文章同行评审

摘要

Online auralization in virtual reality (VR) requires fast yet perceptually accurate synthesis of spatial room impulse responses (SRIRs) under complex geometries, frequency-dependent materials, and dynamically changing source-listener configurations. Existing learning-based methods generally concentrate on generating only the short early-reflection portion of SRIRs, while full-response simulation with geometric acoustics (GA) incurs substantial computational cost at high reflection orders, leading to a trade-off between acoustic fidelity and practical runtime efficiency. In this work, we propose SRERS, a scene-waveform multimodal framework that generates full-length SRIRs from a face-based scene representation with acoustic attributes, source-listener coordinates, and a low-order reflection (LoR) computed via GA. The LoR is incorporated as an auxiliary modality to provide a physically grounded temporal anchor for sparse early arrivals, enabling the network to learn residual components beyond the LoR that model scattering, occlusion, and frequency-dependent coloration. To support training and generalization under diverse acoustic conditions, we further construct a dedicated SRIR dataset with enhanced variability. Experimental results demonstrate that SRERS consistently outperforms state-of-the-art baselines in both full-length and 4096-sample early-reflection SRIR generation. The performance gains remain evident under a strict w/o LoR protocol that excludes deterministic LoR contributions. Complexity analysis shows that SRERS enables online SRIR updating with controllable computational cost, while subjective listening evaluations validate improved perceptual similarity and robustness across conditions.

源语言英语
页(从-至)3854-3869
页数16
期刊IEEE Transactions on Audio, Speech and Language Processing
34
DOI
出版状态已出版 - 2026

学术指纹

探究 'A Multimodal Deep Learning Framework for Spatial Room Impulse Response Generation in VR Auralization' 的科研主题。它们共同构成独一无二的学术指纹。

引用此