跳到主要导航 跳到搜索 跳到主要内容

AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition

  • Yuhang Dai
  • , He Wang
  • , Xingchen Li
  • , Zihan Zhang
  • , Shuiyuan Wang
  • , Lei Xie
  • , Xin Xu
  • , Hongxiao Guo
  • , Shaoji Zhang
  • , Hui Bu
  • , Wei Chen
  • Northwestern Polytechnical University Xian
  • Beijing AISHELL Technology Co., Ltd.
  • Li Auto Inc.

科研成果: 期刊稿件会议文章同行评审

2 引用 (Scopus)

摘要

This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an electric vehicle across more than 60 real driving scenarios. This audio data consists of four far-field speech signals captured by microphones located on each car door, as well as near-field signals obtained from high-fidelity headset microphones worn by each speaker. (2) a collection of 40 hours of real-world environmental noise recordings, which supports the in-car speech data simulation. Moreover, we also provide an open-access, reproducible baseline system based on this dataset. This system features a speech frontend model that employs speech source separation to extract each speaker's clean speech from the far-field signals, along with a speech recognition module that accurately transcribes the content of each individual speaker. Experimental results demonstrate the challenges faced by various mainstream ASR models when evaluated on the AISHELL-5. We firmly believe the AISHELL-5 dataset will significantly advance the research on ASR systems under complex driving scenarios by establishing the first publicly available in-car ASR benchmark.

源语言英语
页(从-至)5493-5497
页数5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
DOI
出版状态已出版 - 2025
活动26th Interspeech Conference 2025 - Rotterdam, 荷兰
期限: 17 8月 202521 8月 2025

指纹

探究 'AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition' 的科研主题。它们共同构成独一无二的指纹。

引用此