跳到主要导航 跳到搜索 跳到主要内容

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

  • Linhan Ma
  • , Dake Guo
  • , Kun Song
  • , Yuepeng Jiang
  • , Shuai Wang
  • , Liumeng Xue
  • , Weiming Xu
  • , Huan Zhao
  • , Binbin Zhang
  • , Lei Xie
  • Northwestern Polytechnical University Xian
  • Shenzhen Research Institute of Big Data
  • The Chinese University of Hong Kong, Shenzhen
  • WeNet Open Source Community

科研成果: 期刊稿件会议文章同行评审

23 引用 (Scopus)

摘要

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus derived from the open-sourced WenetSpeech dataset. Tailored for the text-to-speech tasks, we refined WenetSpeech by adjusting segment boundaries, enhancing the audio quality, and eliminating speaker mixing within each segment. Following a more accurate transcription process and quality-based data filtering process, the obtained WenetSpeech4TTS corpus contains 12, 800 hours of paired audio-text data. Furthermore, we have created subsets of varying sizes, categorized by segment quality scores to allow for TTS model training and fine-tuning. VALLE and NaturalSpeech 2 systems are trained and fine-tuned on these subsets to validate the usability of WenetSpeech4TTS, establishing baselines on benchmark for fair comparison of TTS systems. The corpus and corresponding benchmarks are publicly available on huggingface.

源语言英语
页(从-至)1840-1844
页数5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
DOI
出版状态已出版 - 2024
活动25th Interspeech Conferece 2024 - Kos Island, 希腊
期限: 1 9月 20245 9月 2024

指纹

探究 'WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark' 的科研主题。它们共同构成独一无二的指纹。

引用此