跳到主要导航 跳到搜索 跳到主要内容

Learning acoustic word embeddings with temporal context for query-by-example speech search

  • Yougen Yuan
  • , Cheung Chi Leung
  • , Lei Xie
  • , Hongjie Chen
  • , Bin Ma
  • , Haizhou Li
  • Northwestern Polytechnical University Xian
  • Alibaba Group Holding Ltd.
  • National University of Singapore

科研成果: 期刊稿件会议文章同行评审

33 引用 (Scopus)

摘要

We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search. The temporal context includes the leading and trailing word sequences of a word. We assume that there exist spoken word pairs in the training database. We pad the word pairs with their original temporal context to form fixed-length speech segment pairs. We obtain the acoustic word embeddings through a deep convolutional neural network (CNN) which is trained on the speech segment pairs with a triplet loss. By shifting a fixed-length analysis window through the search content, we obtain a running sequence of embeddings. In this way, searching for the spoken query is equivalent to the matching of acoustic word embeddings. The experiments show that our proposed acoustic word embeddings learned with temporal context are effective in QbE speech search. They outperform the state-of-the-art frame-level feature representations and reduce run-time computation since no dynamic time warping is required in QbE speech search. We also find that it is important to have sufficient speech segment pairs to train the deep CNN for effective acoustic word embeddings.

源语言英语
页(从-至)97-101
页数5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2018-September
DOI
出版状态已出版 - 2018
活动19th Annual Conference of the International Speech Communication, INTERSPEECH 2018 - Hyderabad, 印度
期限: 2 9月 20186 9月 2018

学术指纹

探究 'Learning acoustic word embeddings with temporal context for query-by-example speech search' 的科研主题。它们共同构成独一无二的学术指纹。

引用此