跳到主要导航 跳到搜索 跳到主要内容

Audio-Visual Wake Word Spotting in MISP2021 Challenge: Dataset Release and Deep Analysis

  • Hengshun Zhou
  • , Jun Du
  • , Gongzhen Zou
  • , Zhaoxu Nian
  • , Chin Hui Lee
  • , Sabato Marco Siniscalchi
  • , Shinji Watanabe
  • , Odette Scharenborg
  • , Jingdong Chen
  • , Shifu Xiong
  • , Jian Qing Gao
  • University of Science and Technology of China
  • Georgia Institute of Technology
  • Kore University of Enna
  • Carnegie Mellon University
  • Delft University of Technology
  • IFLYTEK Co., Ltd.

科研成果: 期刊稿件会议文章同行评审

12 引用 (Scopus)

摘要

In this paper, we describe and release publicly the audio-visual wake word spotting (WWS) database in the MISP2021 Challenge, which covers a range of scenarios of audio and video data collected by near-, mid-, and far-field microphone arrays, and cameras, to create a shared and publicly available database for WWS. The database and the code 2 are released, which will be a valuable addition to the community for promoting WWS research using multi-modality information in realistic and complex conditions. Moreover, we investigated the different data augmentation methods for single modalities on an end-to-end WWS network. A set of audio-visual fusion experiments and analysis were conducted to observe the assistance from visual information to acoustic information based on different audio and video field configurations. The results showed that the fusion system generally improves over the single-modality (audio- or video-only) system, especially under complex noisy conditions.

源语言英语
页(从-至)1111-1115
页数5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2022-September
DOI
出版状态已出版 - 2022
活动23rd Annual Conference of the International Speech Communication Association, INTERSPEECH 2022 - Incheon, 韩国
期限: 18 9月 202222 9月 2022

指纹

探究 'Audio-Visual Wake Word Spotting in MISP2021 Challenge: Dataset Release and Deep Analysis' 的科研主题。它们共同构成独一无二的指纹。

引用此