TY - GEN
T1 - Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
AU - Geng, Xuelong
AU - Xu, Tianyi
AU - Wei, Kun
AU - Mu, Bingshen
AU - Xue, Hongfei
AU - Wang, He
AU - Li, Yangze
AU - Guo, Pengcheng
AU - Dai, Yuhang
AU - Li, Longhao
AU - Shao, Mingchen
AU - Xie, Lei
N1 - Publisher Copyright:
©2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm. Building upon this momentum, our research delves into an in-depth examination of this paradigm on a large open-source Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoder-LLM ASR paradigm. Furthermore, we introduce a three-stage training approach, expressly developed to enhance the model’s ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL-1, Test Net, and Test Meeting test sets. Our analysis presents an empirical foundation for future research in LLM-based ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pre-trained models and training logs to promote reproducible research.
AB - Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm. Building upon this momentum, our research delves into an in-depth examination of this paradigm on a large open-source Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoder-LLM ASR paradigm. Furthermore, we introduce a three-stage training approach, expressly developed to enhance the model’s ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL-1, Test Net, and Test Meeting test sets. Our analysis presents an empirical foundation for future research in LLM-based ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pre-trained models and training logs to promote reproducible research.
KW - LLM
KW - speech foundation model
KW - speech recognition
UR - https://www.scopus.com/pages/publications/85205819626
U2 - 10.1109/ISCSLP63861.2024.10800077
DO - 10.1109/ISCSLP63861.2024.10800077
M3 - 会议稿件
AN - SCOPUS:85205819626
T3 - 2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
SP - 26
EP - 30
BT - 2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
A2 - Qian, Yanmin
A2 - Jin, Qin
A2 - Ou, Zhijian
A2 - Ling, Zhenhua
A2 - Wu, Zhiyong
A2 - Li, Ya
A2 - Xie, Lei
A2 - Tao, Jianhua
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
Y2 - 7 November 2024 through 10 November 2024
ER -