Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher

Heyang Xue; Shan Yang; Yi Lei; Lei Xie; Xiulin Li

doi:10.1109/SLT48900.2021.9383585

Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher

Heyang Xue, Shan Yang, Yi Lei, Lei Xie, Xiulin Li

School of Computer Science

Northwestern Polytechnical University Xian

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › peer-review

9 Scopus citations

Abstract

Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produce a natural singing voice from lyrics and music-related transcription. However, such a corpus is difficult to collect since it's hard for many of us to sing like a professional singer. In this paper, we propose an approach - Learn2Sing that only needs a singing teacher to generate the target speakers' singing voice without their singing voice data. In our approach, a teacher's singing corpus and speech from multiple target speakers are trained in a frame-level auto-regressive acoustic model where singing and speaking share the common speaker embedding and style tag embedding. Meanwhile, since there is no music-related transcription for the target speaker, we use log-scale fundamental frequency (LF0) as an auxiliary feature as the inputs of the acoustic model for building a unified input representation. In order to enable the target speaker to sing without singing reference audio in the inference stage, a duration model and an LF0 prediction model are also trained. Particularly, we employ domain adversarial training (DAT) in the acoustic model, which aims to enhance the singing performance of target speakers by disentangling style from acoustic features of singing and speaking data. Our experiments indicate that the proposed approach is capable of synthesizing singing voice for target speaker given only their speech samples.

Original language	English
Title of host publication	2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings
Publisher	Institute of Electrical and Electronics Engineers Inc.
Pages	522-529
Number of pages	8
ISBN (Electronic)	9781728170664
DOIs	https://doi.org/10.1109/SLT48900.2021.9383585
State	Published - 19 Jan 2021
Event	2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Virtual, Shenzhen, China Duration: 19 Jan 2021 → 22 Jan 2021

Publication series

Name	2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings

Conference

Conference	2021 IEEE Spoken Language Technology Workshop, SLT 2021
Country/Territory	China
City	Virtual, Shenzhen
Period	19/01/21 → 22/01/21

Keywords

auto-regressive model
singing voice synthesis
text-to-singing

Access to Document

10.1109/SLT48900.2021.9383585

Cite this

Xue, H., Yang, S., Lei, Y., Xie, L., & Li, X. (2021). Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher. In 2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings (pp. 522-529). Article 9383585 (2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/SLT48900.2021.9383585

@inproceedings{d0cddc7932f6484cb6a0ea19ec9c553e,

title = "Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher",

abstract = "Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produce a natural singing voice from lyrics and music-related transcription. However, such a corpus is difficult to collect since it's hard for many of us to sing like a professional singer. In this paper, we propose an approach - Learn2Sing that only needs a singing teacher to generate the target speakers' singing voice without their singing voice data. In our approach, a teacher's singing corpus and speech from multiple target speakers are trained in a frame-level auto-regressive acoustic model where singing and speaking share the common speaker embedding and style tag embedding. Meanwhile, since there is no music-related transcription for the target speaker, we use log-scale fundamental frequency (LF0) as an auxiliary feature as the inputs of the acoustic model for building a unified input representation. In order to enable the target speaker to sing without singing reference audio in the inference stage, a duration model and an LF0 prediction model are also trained. Particularly, we employ domain adversarial training (DAT) in the acoustic model, which aims to enhance the singing performance of target speakers by disentangling style from acoustic features of singing and speaking data. Our experiments indicate that the proposed approach is capable of synthesizing singing voice for target speaker given only their speech samples.",

keywords = "auto-regressive model, singing voice synthesis, text-to-singing",

author = "Heyang Xue and Shan Yang and Yi Lei and Lei Xie and Xiulin Li",

note = "Publisher Copyright: {\textcopyright} 2021 IEEE.; 2021 IEEE Spoken Language Technology Workshop, SLT 2021 ; Conference date: 19-01-2021 Through 22-01-2021",

year = "2021",

month = jan,

day = "19",

doi = "10.1109/SLT48900.2021.9383585",

language = "英语",

series = "2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

pages = "522--529",

booktitle = "2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings",

}

Xue, H, Yang, S, Lei, Y, Xie, L & Li, X 2021, Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher. in 2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings., 9383585, 2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings, Institute of Electrical and Electronics Engineers Inc., pp. 522-529, 2021 IEEE Spoken Language Technology Workshop, SLT 2021, Virtual, Shenzhen, China, 19/01/21. https://doi.org/10.1109/SLT48900.2021.9383585

Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher. / Xue, Heyang; Yang, Shan; Lei, Yi et al.
2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings. Institute of Electrical and Electronics Engineers Inc., 2021. p. 522-529 9383585 (2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings).

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › peer-review

TY - GEN

T1 - Learn2Sing

T2 - 2021 IEEE Spoken Language Technology Workshop, SLT 2021

AU - Xue, Heyang

AU - Yang, Shan

AU - Lei, Yi

AU - Xie, Lei

AU - Li, Xiulin

PY - 2021/1/19

Y1 - 2021/1/19

N2 - Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produce a natural singing voice from lyrics and music-related transcription. However, such a corpus is difficult to collect since it's hard for many of us to sing like a professional singer. In this paper, we propose an approach - Learn2Sing that only needs a singing teacher to generate the target speakers' singing voice without their singing voice data. In our approach, a teacher's singing corpus and speech from multiple target speakers are trained in a frame-level auto-regressive acoustic model where singing and speaking share the common speaker embedding and style tag embedding. Meanwhile, since there is no music-related transcription for the target speaker, we use log-scale fundamental frequency (LF0) as an auxiliary feature as the inputs of the acoustic model for building a unified input representation. In order to enable the target speaker to sing without singing reference audio in the inference stage, a duration model and an LF0 prediction model are also trained. Particularly, we employ domain adversarial training (DAT) in the acoustic model, which aims to enhance the singing performance of target speakers by disentangling style from acoustic features of singing and speaking data. Our experiments indicate that the proposed approach is capable of synthesizing singing voice for target speaker given only their speech samples.

AB - Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produce a natural singing voice from lyrics and music-related transcription. However, such a corpus is difficult to collect since it's hard for many of us to sing like a professional singer. In this paper, we propose an approach - Learn2Sing that only needs a singing teacher to generate the target speakers' singing voice without their singing voice data. In our approach, a teacher's singing corpus and speech from multiple target speakers are trained in a frame-level auto-regressive acoustic model where singing and speaking share the common speaker embedding and style tag embedding. Meanwhile, since there is no music-related transcription for the target speaker, we use log-scale fundamental frequency (LF0) as an auxiliary feature as the inputs of the acoustic model for building a unified input representation. In order to enable the target speaker to sing without singing reference audio in the inference stage, a duration model and an LF0 prediction model are also trained. Particularly, we employ domain adversarial training (DAT) in the acoustic model, which aims to enhance the singing performance of target speakers by disentangling style from acoustic features of singing and speaking data. Our experiments indicate that the proposed approach is capable of synthesizing singing voice for target speaker given only their speech samples.

KW - auto-regressive model

KW - singing voice synthesis

KW - text-to-singing

UR - http://www.scopus.com/inward/record.url?scp=85103987827&partnerID=8YFLogxK

U2 - 10.1109/SLT48900.2021.9383585

DO - 10.1109/SLT48900.2021.9383585

M3 - 会议稿件

AN - SCOPUS:85103987827

T3 - 2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings

SP - 522

EP - 529

BT - 2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings

PB - Institute of Electrical and Electronics Engineers Inc.

Y2 - 19 January 2021 through 22 January 2021

ER -

Xue H, Yang S, Lei Y, Xie L, Li X. Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher. In 2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings. Institute of Electrical and Electronics Engineers Inc. 2021. p. 522-529. 9383585. (2021 IEEE Spoken Language Technology Workshop, SLT 2021 - Proceedings). doi: 10.1109/SLT48900.2021.9383585

Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing Teacher

Abstract

Publication series

Conference

Keywords

Access to Document

Other files and links

Fingerprint

Cite this