Skip to main navigation Skip to search Skip to main content

A speech enhancement system for automotive speech recognition with a hybrid voice activity detection method

  • University of Science and Technology of China

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

10 Scopus citations

Abstract

This paper presents a front-end speech enhancement approach to robust speech recognition in automotive environments. It combines hybrid voice activity detection (VAD), relative transfer function (RT-F) based generalized sidelobe cancelation, and single-channel post filtering to enhance the speech signal of interest, thereby improving the robustness of speech recognition. First, we choose four typical driving scenarios, which include most of the noise types in automobiles to record training data. The recorded data is then used to train deep neural network models (DNNs) for both speech and noise. The trained DNNs are subsequently used to estimate the speech presence probability on a frame-by-frame basis. This speech presence probability is then combined with the output of an energy-based VAD to form a hybrid VAD, which serves as the basis for the rest components of the speech enhancement system, including RTF estimation, adaptive beamforming, and post-filtering. Experiments are conducted in real automotive environments. The results show that the developed method can significantly improve the performance of both VAD and automatic speech recognition (ASR).

Original languageEnglish
Title of host publication16th International Workshop on Acoustic Signal Enhancement, IWAENC 2018 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages456-460
Number of pages5
ISBN (Electronic)9781538681510
DOIs
StatePublished - 2 Nov 2018
Event16th International Workshop on Acoustic Signal Enhancement, IWAENC 2018 - Tokyo, Japan
Duration: 17 Sep 201820 Sep 2018

Publication series

Name16th International Workshop on Acoustic Signal Enhancement, IWAENC 2018 - Proceedings

Conference

Conference16th International Workshop on Acoustic Signal Enhancement, IWAENC 2018
Country/TerritoryJapan
CityTokyo
Period17/09/1820/09/18

Keywords

  • Deep neural network
  • Microphone array
  • Speech enhancement
  • Speech recognition
  • Voice activity detection

Fingerprint

Dive into the research topics of 'A speech enhancement system for automotive speech recognition with a hybrid voice activity detection method'. Together they form a unique fingerprint.

Cite this