Abstract
Online auralization in virtual reality (VR) requires fast yet perceptually accurate synthesis of spatial room impulse responses (SRIRs) under complex geometries, frequency-dependent materials, and dynamically changing source-listener configurations. Existing learning-based methods generally concentrate on generating only the short early-reflection portion of SRIRs, while full-response simulation with geometric acoustics (GA) incurs substantial computational cost at high reflection orders, leading to a trade-off between acoustic fidelity and practical runtime efficiency. In this work, we propose SRERS, a scene-waveform multimodal framework that generates full-length SRIRs from a face-based scene representation with acoustic attributes, source-listener coordinates, and a low-order reflection (LoR) computed via GA. The LoR is incorporated as an auxiliary modality to provide a physically grounded temporal anchor for sparse early arrivals, enabling the network to learn residual components beyond the LoR that model scattering, occlusion, and frequency-dependent coloration. To support training and generalization under diverse acoustic conditions, we further construct a dedicated SRIR dataset with enhanced variability. Experimental results demonstrate that SRERS consistently outperforms state-of-the-art baselines in both full-length and 4096-sample early-reflection SRIR generation. The performance gains remain evident under a strict w/o LoR protocol that excludes deterministic LoR contributions. Complexity analysis shows that SRERS enables online SRIR updating with controllable computational cost, while subjective listening evaluations validate improved perceptual similarity and robustness across conditions.
| Original language | English |
|---|---|
| Pages (from-to) | 3854-3869 |
| Number of pages | 16 |
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 34 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Spatial room impulse response
- auralization
- multimodal deep learning
- virtual reality
Fingerprint
Dive into the research topics of 'A Multimodal Deep Learning Framework for Spatial Room Impulse Response Generation in VR Auralization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver