Skip to main navigation Skip to search Skip to main content

DySiME: Dynamic Single-Source Multichannel Enhancement Using Time-Varying Directional Cues

  • Waseda University
  • Northwestern Polytechnical University Xian

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Single-source multichannel speech enhancement involves extracting speech from a desired, target speaker in noisy, reverberant environments while maintaining spatial fidelity across multichannel outputs. A recent solution to this problem, multichannel-to-multichannel target sound extraction (M2MTSE), utilizes directional and temporal cues to perform end-toend complex spectrogram mapping in reverberant environments with static sources. However, its effectiveness is limited by heavy reliance on prior knowledge and handcrafted cyclic positional embeddings, reducing its practicality in real-world applications. To overcome these limitations in dynamic source scenarios where the target speaker is moving, we propose DySiME, dynamic single-source multichannel enhancement, an end-to-end framework tailored for moving sources. DySiME integrates a direction-of-arrival (DOA) estimation module based on full-band and narrow-band fusion for sound source localization (FN-SSL) to continuously track the target source direction. Furthermore, a learnable positional-information adapter incorporates intermediate features from the DOA estimator into the enhancement backbone, enabling the model to utilize time-varying spatial cues for more effective speech enhancement. This design reduces reliance on prior DOA knowledge during inference and enables robust, real-time enhancement of moving sources. We evaluate the system using a simulated 4 -channel circular microphone array, and the results show that DySiME consistently outperforms the baseline in both speech quality and spatial accuracy.

Original languageEnglish
Title of host publication2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages154-159
Number of pages6
ISBN (Electronic)9798331572068
DOIs
StatePublished - 2025
Event17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025 - Singapore, Singapore
Duration: 22 Oct 202524 Oct 2025

Publication series

Name2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025

Conference

Conference17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Country/TerritorySingapore
CitySingapore
Period22/10/2524/10/25

Fingerprint

Dive into the research topics of 'DySiME: Dynamic Single-Source Multichannel Enhancement Using Time-Varying Directional Cues'. Together they form a unique fingerprint.

Cite this