跳到主要导航 跳到搜索 跳到主要内容

Promptspeaker: Speaker Generation Based on Text Descriptions

  • Yongmao Zhang
  • , Guanghou Liu
  • , Yi Lei
  • , Yunlin Chen
  • , Hao Yin
  • , Lei Xie
  • , Zhifei Li
  • Northwestern Polytechnical University Xian
  • Ltd

科研成果: 书/报告/会议事项章节会议稿件同行评审

14 引用 (Scopus)

摘要

Recently, text-guided content generation has received extensive attention. In this work, we explore the possibility of text description-based speaker generation, i.e., using text prompts to control the speaker generation process. Specifically, we propose PromptSpeaker, a text-guided speaker generation system. PromptSpeaker consists of a prompt encoder, a zero-shot VITS, and a Glow model, where the prompt encoder predicts a prior distribution based on the text description and samples from this distribution to obtain a semantic representation. The Glow model subsequently converts the semantic representation into a speaker representation, and the zero-shot VITS finally synthesizes the speaker's voice based on the speaker representation. We verify that PromptSpeaker can generate speakers new from the training set by objective metrics, and the synthetic speaker voice has reasonable subjective matching quality with the speaker prompt. Our audio samples are available on the demo website11Demo: https://promptspeaker.github.io/demo/

源语言英语
主期刊名2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798350306897
DOI
出版状态已出版 - 2023
活动2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023 - Taipei, 中国台湾
期限: 16 12月 202320 12月 2023

丛书

姓名2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023

会议

会议2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023
国家/地区中国台湾
Taipei
时期16/12/2320/12/23

学术指纹

探究 'Promptspeaker: Speaker Generation Based on Text Descriptions' 的科研主题。它们共同构成独一无二的学术指纹。

引用此