跳到主要导航 跳到搜索 跳到主要内容

Simplified Self-Attention for Transformer-Based end-to-end Speech Recognition

  • Haoneng Luo
  • , Shiliang Zhang
  • , Ming Lei
  • , Lei Xie
  • Northwestern Polytechnical University Xian

科研成果: 书/报告/会议事项章节会议稿件同行评审

31 引用 (Scopus)

摘要

Transformer models have been introduced into end-to-end speech recognition with state-of-the-art performance on various tasks owing to their superiority in modeling long-term dependencies. However, such improvements are usually obtained through the use of very large neural networks. Transformer models mainly include two submodules - position-wise feedforward layers and self-attention (SAN) layers. In this paper, to reduce the model complexity while maintaining good performance, we propose a simplified self-attention (SSAN) layer which employs FSMN memory blocks instead of projection layers to form query and key vectors for transformer-based end-to-end speech recognition. We evaluate the SSAN-based and the conventional SAN-based transformers on the public AISHELL-1, internal 1000-hour and 20,000-hour large-scale Mandarin tasks. Results show that our proposed SSAN-based transformer model can achieve over 20% reduction in model parameters and 6.7% relative CER reduction on the AISHELL-1 task. With impressively 20% parameter reduction, our model shows no loss of recognition performance on the 20,000-hour large-scale task.

源语言英语
主期刊名2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings
出版商Institute of Electrical and Electronics Engineers Inc.
75-81
页数7
ISBN(电子版)9781728170664
DOI
出版状态已出版 - 19 1月 2021
活动2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Virtual, Online, 中国
期限: 19 1月 202122 1月 2021

出版系列

姓名2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings

会议

会议2021 IEEE Spoken Language Technology Workshop, SLT 2021
国家/地区中国
Virtual, Online
时期19/01/2122/01/21

学术指纹

探究 'Simplified Self-Attention for Transformer-Based end-to-end Speech Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此