摘要
Emotional voice conversion aims at converting speech from one emotion state to another. This paper proposes to model tim- bre and prosody features using a deep bidirectional long short- term memory (DBLSTM) for emotional voice conversion. A continuous wavelet transform (CWT) representation of funda- mental frequency (F0) and energy contour are used for prosody modeling. Specifically, we use CWT to decompose F0 into a five-scale representation, and decompose energy contour into a ten-scale representation, where each feature scale corresponds to a temporal scale. Both spectrum and prosody (F0 and energy contour) features are simultaneously converted by a sequence to sequence conversion method with DBLSTM model, which captures both frame-wise and long-range relationship between source and target voice. The converted speech signals are e- valuated both objectively and subjectively, which confirms the effectiveness of the proposed method.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 2453-2457 |
| 页数 | 5 |
| 期刊 | Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH |
| 卷 | 08-12-September-2016 |
| DOI | |
| 出版状态 | 已出版 - 2016 |
| 活动 | 17th Annual Conference of the International Speech Communication Association, INTERSPEECH 2016 - San Francisco, 美国 期限: 8 9月 2016 → 16 9月 2016 |
学术指纹
探究 'Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver