摘要
The recently proposed serialized output training (SOT) simplifies multi-talker automatic speech recognition (ASR) by generating speaker transcriptions separated by a special token. However, frequent speaker changes can make speaker change prediction difficult. To address this, we propose boundary-aware serialized output training (BA-SOT), which explicitly incorporates boundary knowledge into the decoder via a speaker change detection task and boundary constraint loss. We also introduce a two-stage connectionist temporal classification (CTC) strategy that incorporates token-level SOT CTC to restore temporal context information. Besides typical character error rate (CER), we introduce utterance-dependent character error rate (UD-CER) to further measure the precision of speaker change prediction. Compared to original SOT, BA-SOT reduces CER/UD-CER by 5.1%/14.0%, and leveraging a pre-trained ASR model for BA-SOT model initialization further reduces CER/UD-CER by 8.4%/19.9%.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 3487-3491 |
| 页数 | 5 |
| 期刊 | Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH |
| 卷 | 2023-August |
| DOI | |
| 出版状态 | 已出版 - 2023 |
| 活动 | 24th Annual conference of the International Speech Communication Association, Interspeech 2023 - Dublin, 爱尔兰 期限: 20 8月 2023 → 24 8月 2023 |
学术指纹
探究 'BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver