跳到主要导航 跳到搜索 跳到主要内容

S-DCCRN: SUPER WIDE BAND DCCRN WITH LEARNABLE COMPLEX FEATURE FOR SPEECH ENHANCEMENT

  • Shubo Lv
  • , Yihui Fu
  • , Mengtao Xing
  • , Jiayao Sun
  • , Lei Xie
  • , Jun Huang
  • , Yannan Wang
  • , Tao Yu
  • Northwestern Polytechnical University Xian
  • Tencent

科研成果: 书/报告/会议事项章节会议稿件同行评审

53 引用 (Scopus)

摘要

In speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band signal with a sampling rate of 16K Hz. However, research on super wide band (e.g., 32K Hz) or even full-band (48K) denoising using deep learning is still in its infancy due to the difficulty of modeling more frequency bands and particularly high frequency components. In this paper, we extend our previous deep complex convolution recurrent neural network (DCCRN) substantially to a super wide band version - S-DCCRN, to perform speech denoising on speech of 32K Hz sampling rate. We first employ a cascaded sub-band and full-band processing module, which consists of two small-footprint DCCRNs - one operates on sub-band signal and one operates on full-band signal, aiming at benefiting from both local and global frequency information. Moreover, instead of simply adopting the STFT feature as input, we use a complex feature encoder trained in an end-to-end manner to refine the information of different frequency bands. We also use a complex feature decoder to revert the feature to time-frequency domain. Finally, a learnable spectrum compression method is adopted to adjust the energy of different frequency bands, which is beneficial for neural network learning. The proposed model, S-DCCRN, has surpassed PercepNet as well as several competitive models and achieves state-of-the-art performance in terms of speech quality and intelligibility. Ablation studies also demonstrate the effectiveness of different contributions.

源语言英语
主期刊名2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings
出版商Institute of Electrical and Electronics Engineers Inc.
7767-7771
页数5
ISBN(电子版)9781665405409
DOI
出版状态已出版 - 2022
活动2022 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022 - Hybrid, 新加坡
期限: 22 5月 202227 5月 2022

出版系列

姓名ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
2022-May
ISSN(印刷版)1520-6149

会议

会议2022 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022
国家/地区新加坡
Hybrid
时期22/05/2227/05/22

指纹

探究 'S-DCCRN: SUPER WIDE BAND DCCRN WITH LEARNABLE COMPLEX FEATURE FOR SPEECH ENHANCEMENT' 的科研主题。它们共同构成独一无二的指纹。

引用此