TY - GEN
T1 - DySiME
T2 - 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
AU - Liang, Hao
AU - Yang, Yichen
AU - Zhang, Xiao
AU - Makino, Shoji
AU - Chen, Jingdong
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Single-source multichannel speech enhancement involves extracting speech from a desired, target speaker in noisy, reverberant environments while maintaining spatial fidelity across multichannel outputs. A recent solution to this problem, multichannel-to-multichannel target sound extraction (M2MTSE), utilizes directional and temporal cues to perform end-toend complex spectrogram mapping in reverberant environments with static sources. However, its effectiveness is limited by heavy reliance on prior knowledge and handcrafted cyclic positional embeddings, reducing its practicality in real-world applications. To overcome these limitations in dynamic source scenarios where the target speaker is moving, we propose DySiME, dynamic single-source multichannel enhancement, an end-to-end framework tailored for moving sources. DySiME integrates a direction-of-arrival (DOA) estimation module based on full-band and narrow-band fusion for sound source localization (FN-SSL) to continuously track the target source direction. Furthermore, a learnable positional-information adapter incorporates intermediate features from the DOA estimator into the enhancement backbone, enabling the model to utilize time-varying spatial cues for more effective speech enhancement. This design reduces reliance on prior DOA knowledge during inference and enables robust, real-time enhancement of moving sources. We evaluate the system using a simulated 4 -channel circular microphone array, and the results show that DySiME consistently outperforms the baseline in both speech quality and spatial accuracy.
AB - Single-source multichannel speech enhancement involves extracting speech from a desired, target speaker in noisy, reverberant environments while maintaining spatial fidelity across multichannel outputs. A recent solution to this problem, multichannel-to-multichannel target sound extraction (M2MTSE), utilizes directional and temporal cues to perform end-toend complex spectrogram mapping in reverberant environments with static sources. However, its effectiveness is limited by heavy reliance on prior knowledge and handcrafted cyclic positional embeddings, reducing its practicality in real-world applications. To overcome these limitations in dynamic source scenarios where the target speaker is moving, we propose DySiME, dynamic single-source multichannel enhancement, an end-to-end framework tailored for moving sources. DySiME integrates a direction-of-arrival (DOA) estimation module based on full-band and narrow-band fusion for sound source localization (FN-SSL) to continuously track the target source direction. Furthermore, a learnable positional-information adapter incorporates intermediate features from the DOA estimator into the enhancement backbone, enabling the model to utilize time-varying spatial cues for more effective speech enhancement. This design reduces reliance on prior DOA knowledge during inference and enables robust, real-time enhancement of moving sources. We evaluate the system using a simulated 4 -channel circular microphone array, and the results show that DySiME consistently outperforms the baseline in both speech quality and spatial accuracy.
UR - https://www.scopus.com/pages/publications/105030454253
U2 - 10.1109/APSIPAASC65261.2025.11249356
DO - 10.1109/APSIPAASC65261.2025.11249356
M3 - 会议稿件
AN - SCOPUS:105030454253
T3 - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
SP - 154
EP - 159
BT - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 22 October 2025 through 24 October 2025
ER -