Skip to main navigation Skip to search Skip to main content

Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets

  • Xuelong Geng
  • , Tianyi Xu
  • , Kun Wei
  • , Bingshen Mu
  • , Hongfei Xue
  • , He Wang
  • , Yangze Li
  • , Pengcheng Guo
  • , Yuhang Dai
  • , Longhao Li
  • , Mingchen Shao
  • , Lei Xie
  • Northwestern Polytechnical University Xian

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

10 Scopus citations

Abstract

Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm. Building upon this momentum, our research delves into an in-depth examination of this paradigm on a large open-source Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoder-LLM ASR paradigm. Furthermore, we introduce a three-stage training approach, expressly developed to enhance the model’s ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL-1, Test Net, and Test Meeting test sets. Our analysis presents an empirical foundation for future research in LLM-based ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pre-trained models and training logs to promote reproducible research.

Original languageEnglish
Title of host publication2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
EditorsYanmin Qian, Qin Jin, Zhijian Ou, Zhenhua Ling, Zhiyong Wu, Ya Li, Lei Xie, Jianhua Tao
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages26-30
Number of pages5
ISBN (Electronic)9798331516826
DOIs
StatePublished - 2024
Event14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024 - Beijing, China
Duration: 7 Nov 202410 Nov 2024

Publication series

Name2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024

Conference

Conference14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
Country/TerritoryChina
CityBeijing
Period7/11/2410/11/24

Keywords

  • LLM
  • speech foundation model
  • speech recognition

Fingerprint

Dive into the research topics of 'Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets'. Together they form a unique fingerprint.

Cite this