跳到主要导航 跳到搜索 跳到主要内容

An end-to-end architecture of online multi-channel speech separation

  • Jian Wu
  • , Zhuo Chen
  • , Jinyu Li
  • , Takuya Yoshioka
  • , Zhili Tan
  • , Ed Lin
  • , Yi Luo
  • , Lei Xie
  • Northwestern Polytechnical University Xian
  • Microsoft USA

科研成果: 书/报告/会议事项章节会议稿件同行评审

21 引用 (Scopus)

摘要

Multi-speaker speech recognition has been one of the key challenges in conversation transcription as it breaks the single active speaker assumption employed by most state-of-the-art speech recognition systems. Speech separation is considered as a remedy to this problem. Previously, we introduced a system, called unmixing, fixed-beamformer and extraction (UFE), that was shown to be effective in addressing the speech overlap problem in conversation transcription. With UFE, an input mixed signal is processed by fixed beamformers, followed by a neural network post filtering. Although promising results were obtained, the system contains multiple individually developed modules, leading potentially sub-optimum performance. In this work, we introduce an end-to-end modeling version of UFE. To enable gradient propagation all the way, an attentional selection module is proposed, where an attentional weight is learnt for each beamformer and spatial feature sampled over space. Experimental results show that the proposed system achieves comparable performance in an offline evaluation with the original separate processing-based pipeline, while producing remarkable improvements in an online evaluation.

源语言英语
主期刊名Interspeech 2020
出版商International Speech Communication Association
81-85
页数5
ISBN(印刷版)9781713820697
DOI
出版状态已出版 - 2020
活动21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020 - Shanghai, 中国
期限: 25 10月 202029 10月 2020

出版系列

姓名Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2020-October
ISSN(印刷版)2308-457X
ISSN(电子版)1990-9772

会议

会议21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020
国家/地区中国
Shanghai
时期25/10/2029/10/20

学术指纹

探究 'An end-to-end architecture of online multi-channel speech separation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此