Disentangling Style and Speaker Attributes for TTS Style Transfer

Xiaochun An; Frank K. Soong; Lei Xie

doi:10.1109/TASLP.2022.3145297

Disentangling Style and Speaker Attributes for TTS Style Transfer

Xiaochun An, Frank K. Soong, Lei Xie

School of Computer Science

Research output: Contribution to journal › Article › peer-review

15 Scopus citations

Abstract

End-to-end neural TTS has shown improved performance in speech style transfer. However, the improvement is still limited by the available training data in both target styles and speakers. Additionally, degenerated performance is observed when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style. In this paper, we propose a new approach to seen and unseen style transfer training on disjoint, multi-style datasets, i. e., datasets of different styles are recorded, one individual style by one speaker in multiple utterances. An inverse autoregressive flow (IAF) technique is first introduced to improve the variational inference for learning an expressive style representation. A speaker encoder network is then developed for learning a discriminative speaker embedding, which is jointly trained with the rest neural TTS modules. The proposed approach of seen and unseen style transfer is effectively trained with six specifically-designed objectives: reconstruction loss, adversarial loss, style distortion loss, cycle consistency loss, style classification loss, and speaker classification loss. Experiments demonstrate, both objectively and subjectively, the effectiveness of the proposed approach for seen and unseen style transfer tasks. The performance of our approach is superior to and more robust than those of four other reference systems of prior art.

Original language	English
Pages (from-to)	646-658
Number of pages	13
Journal	IEEE/ACM Transactions on Audio Speech and Language Processing
Volume	30
DOIs	https://doi.org/10.1109/TASLP.2022.3145297
State	Published - 2022

Keywords

Disjoint datasets
neural TTS
style and speaker attributes
style transfer
variational inference

Access to Document

10.1109/TASLP.2022.3145297

Cite this

@article{d42feb53f4454520ab6e92f5b5568ef1,

title = "Disentangling Style and Speaker Attributes for TTS Style Transfer",

abstract = "End-to-end neural TTS has shown improved performance in speech style transfer. However, the improvement is still limited by the available training data in both target styles and speakers. Additionally, degenerated performance is observed when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style. In this paper, we propose a new approach to seen and unseen style transfer training on disjoint, multi-style datasets, i. e., datasets of different styles are recorded, one individual style by one speaker in multiple utterances. An inverse autoregressive flow (IAF) technique is first introduced to improve the variational inference for learning an expressive style representation. A speaker encoder network is then developed for learning a discriminative speaker embedding, which is jointly trained with the rest neural TTS modules. The proposed approach of seen and unseen style transfer is effectively trained with six specifically-designed objectives: reconstruction loss, adversarial loss, style distortion loss, cycle consistency loss, style classification loss, and speaker classification loss. Experiments demonstrate, both objectively and subjectively, the effectiveness of the proposed approach for seen and unseen style transfer tasks. The performance of our approach is superior to and more robust than those of four other reference systems of prior art.",

keywords = "Disjoint datasets, neural TTS, style and speaker attributes, style transfer, variational inference",

author = "Xiaochun An and Soong, {Frank K.} and Lei Xie",

note = "Publisher Copyright: {\textcopyright} 2014 IEEE.",

year = "2022",

doi = "10.1109/TASLP.2022.3145297",

language = "英语",

volume = "30",

pages = "646--658",

journal = "IEEE/ACM Transactions on Audio Speech and Language Processing",

issn = "2329-9290",

publisher = "IEEE Advancing Technology for Humanity",

}

TY - JOUR

T1 - Disentangling Style and Speaker Attributes for TTS Style Transfer

AU - An, Xiaochun

AU - Soong, Frank K.

AU - Xie, Lei

PY - 2022

Y1 - 2022

N2 - End-to-end neural TTS has shown improved performance in speech style transfer. However, the improvement is still limited by the available training data in both target styles and speakers. Additionally, degenerated performance is observed when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style. In this paper, we propose a new approach to seen and unseen style transfer training on disjoint, multi-style datasets, i. e., datasets of different styles are recorded, one individual style by one speaker in multiple utterances. An inverse autoregressive flow (IAF) technique is first introduced to improve the variational inference for learning an expressive style representation. A speaker encoder network is then developed for learning a discriminative speaker embedding, which is jointly trained with the rest neural TTS modules. The proposed approach of seen and unseen style transfer is effectively trained with six specifically-designed objectives: reconstruction loss, adversarial loss, style distortion loss, cycle consistency loss, style classification loss, and speaker classification loss. Experiments demonstrate, both objectively and subjectively, the effectiveness of the proposed approach for seen and unseen style transfer tasks. The performance of our approach is superior to and more robust than those of four other reference systems of prior art.

AB - End-to-end neural TTS has shown improved performance in speech style transfer. However, the improvement is still limited by the available training data in both target styles and speakers. Additionally, degenerated performance is observed when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style. In this paper, we propose a new approach to seen and unseen style transfer training on disjoint, multi-style datasets, i. e., datasets of different styles are recorded, one individual style by one speaker in multiple utterances. An inverse autoregressive flow (IAF) technique is first introduced to improve the variational inference for learning an expressive style representation. A speaker encoder network is then developed for learning a discriminative speaker embedding, which is jointly trained with the rest neural TTS modules. The proposed approach of seen and unseen style transfer is effectively trained with six specifically-designed objectives: reconstruction loss, adversarial loss, style distortion loss, cycle consistency loss, style classification loss, and speaker classification loss. Experiments demonstrate, both objectively and subjectively, the effectiveness of the proposed approach for seen and unseen style transfer tasks. The performance of our approach is superior to and more robust than those of four other reference systems of prior art.

KW - Disjoint datasets

KW - neural TTS

KW - style and speaker attributes

KW - style transfer

KW - variational inference

UR - http://www.scopus.com/inward/record.url?scp=85123765161&partnerID=8YFLogxK

U2 - 10.1109/TASLP.2022.3145297

DO - 10.1109/TASLP.2022.3145297

M3 - 文章

AN - SCOPUS:85123765161

SN - 2329-9290

VL - 30

SP - 646

EP - 658

JO - IEEE/ACM Transactions on Audio Speech and Language Processing

JF - IEEE/ACM Transactions on Audio Speech and Language Processing

ER -

Disentangling Style and Speaker Attributes for TTS Style Transfer

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this