Skip to main navigation Skip to search Skip to main content

Dialospeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching

  • Hanke Xie
  • , Dake Guo
  • , Chengyou Wang
  • , Yue Li
  • , Wenjie Tian
  • , Xinfa Zhu
  • , Xinsheng Wang
  • , Xiulin Li
  • , Guanqiong Miao
  • , Bo Liu
  • , Lei Xie
  • Northwestern Polytechnical University Xian
  • DataBaker (Qingdao) Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However, generating human-like, interactive dialogue speech remains challenging. Current systems face limitations due to the scarcity of dual-track data and difficulties in achieving naturalness, contextual coherence, and interactional dynamics, such as turntaking, overlapping speech, and speaker consistency, in multiturn conversations. To address these challenges, we propose DialoSpeech 11Codes and checkpoints will be publicly released., a dual-track architecture combining a large language model with Chunked Flow Matching for expressive, humanlike dialogue speech synthesis. DialoSpeech generates natural multi-turn conversations with coherent speaker turns and natural overlaps, supporting both Chinese and English and crosslingual speech synthesis. We introduce a data processing pipeline to construct dual-track dialogue datasets, facilitating scalable training and experimental validation. Experiments show that our model outperforms baselines, offering a solution for generating human-like spoken dialogues. Audio samples are available at https://tiamojames.github.io/DialoSpeech/

Original languageEnglish
Title of host publication2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages807-812
Number of pages6
ISBN (Electronic)9798331572068
DOIs
StatePublished - 2025
Event17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025 - Singapore, Singapore
Duration: 22 Oct 202524 Oct 2025

Publication series

Name2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025

Conference

Conference17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Country/TerritorySingapore
CitySingapore
Period22/10/2524/10/25

Keywords

  • Dialogue Generation
  • Flow Matching
  • Language Models

Fingerprint

Dive into the research topics of 'Dialospeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching'. Together they form a unique fingerprint.

Cite this