Enriching source style transfer in recognition-synthesis based non-parallel voice conversion

Zhichao Wang, Xinyong Zhou, Fengyu Yang, Tao Li, Hongqiang Du, Lei Xie, Wendong Gan, Haitao Chen, Hai Li

科研成果: 书/报告/会议事项章节会议稿件同行评审

6 引用 (Scopus)

摘要

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech. This study proposes a source style transfer method based on recognitionsynthesis framework. Previously in speech generation task, prosody can be modeled explicitly with prosodic features or implicitly with a latent prosody extractor. In this paper, taking advantages of both, we model the prosody in a hybrid manner, which effectively combines explicit and implicit methods in a proposed prosody module. Specifically, prosodic features are used to explicit model prosody, while VAE and reference encoder are used to implicitly model prosody, which take Mel spectrum and bottleneck feature as input respectively. Furthermore, adversarial training is introduced to remove speakerrelated information from the VAE outputs, avoiding leaking source speaker information while transferring style. Finally, we use a modified self-attention based encoder to extract sentential context from bottleneck features, which also implicitly aggregates the prosodic aspects of source speech from the layered representations. Experiments show that our approach is superior to the baseline and a competitive system in terms of style transfer; meanwhile, the speech quality and speaker similarity are well maintained.

源语言英语
主期刊名22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021
出版商International Speech Communication Association
4820-4824
页数5
ISBN(电子版)9781713836902
DOI
出版状态已出版 - 2021
活动22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021 - Brno, 捷克共和国
期限: 30 8月 20213 9月 2021

出版系列

姓名Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
6
ISSN(印刷版)2308-457X
ISSN(电子版)1990-9772

会议

会议22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021
国家/地区捷克共和国
Brno
时期30/08/213/09/21

指纹

探究 'Enriching source style transfer in recognition-synthesis based non-parallel voice conversion' 的科研主题。它们共同构成独一无二的指纹。

引用此