TY - JOUR
T1 - Cross-Speaker Emotion Transfer Through Information Perturbation in Emotional Speech Synthesis
AU - Lei, Yi
AU - Yang, Shan
AU - Zhu, Xinfa
AU - Xie, Lei
AU - Su, Dan
N1 - Publisher Copyright:
© 1994-2012 IEEE.
PY - 2022
Y1 - 2022
N2 - Through borrowing emotional expressions from an emotional speaker, cross-speaker emotion transfer is an effective way to produce emotional speech for target speakers without emotional training data. Since emotion and timbre of the source speaker are heavily entangled in speech, existing approaches often struggle to trade off between speaker similarity and emotional expression in the synthetic speech of the target speaker. In this letter, we propose to disentangle timbre and emotion through information perturbation to conduct cross-speaker emotion transfer, which effectively learns the emotional expression of the source speaker and maintains the timbre of the target speaker. Specifically, we separately perturb the timbre and emotion-related features (e.g., formant and pitch) of source speech to obtain and model the timbre- and emotion-independent signals, based on which the proposed model can deliver the emotional expression for target speakers. Experimental results demonstrate the proposed approach significantly outperforms the baselines in terms of naturalness and similarity, indicating the effectiveness of information perturbation for cross-speaker emotion transfer.
AB - Through borrowing emotional expressions from an emotional speaker, cross-speaker emotion transfer is an effective way to produce emotional speech for target speakers without emotional training data. Since emotion and timbre of the source speaker are heavily entangled in speech, existing approaches often struggle to trade off between speaker similarity and emotional expression in the synthetic speech of the target speaker. In this letter, we propose to disentangle timbre and emotion through information perturbation to conduct cross-speaker emotion transfer, which effectively learns the emotional expression of the source speaker and maintains the timbre of the target speaker. Specifically, we separately perturb the timbre and emotion-related features (e.g., formant and pitch) of source speech to obtain and model the timbre- and emotion-independent signals, based on which the proposed model can deliver the emotional expression for target speakers. Experimental results demonstrate the proposed approach significantly outperforms the baselines in terms of naturalness and similarity, indicating the effectiveness of information perturbation for cross-speaker emotion transfer.
KW - Cross-speaker emotion transfer
KW - emotional TTS
KW - information perturbation
KW - speech synthesis
UR - https://www.scopus.com/pages/publications/85137915056
U2 - 10.1109/LSP.2022.3203888
DO - 10.1109/LSP.2022.3203888
M3 - 文章
AN - SCOPUS:85137915056
SN - 1070-9908
VL - 29
SP - 1948
EP - 1952
JO - IEEE Signal Processing Letters
JF - IEEE Signal Processing Letters
ER -