Skip to main navigation Skip to search Skip to main content

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

  • Wenjie Tian
  • , Xinfa Zhu
  • , Haohe Liu
  • , Zhixian Zhao
  • , Zihao Chen
  • , Chaofan Ding
  • , Xinhan Di
  • , Junjie Zheng
  • , Lei Xie
  • Northwestern Polytechnical University Xian
  • University of Surrey
  • Giant Network Inc.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST) generation, which aims to jointly produce synchronized background audio and speech within a unified framework. To tackle V2ST, we introduce DualDub, a unified framework built on a multimodal language model that integrates a multimodal encoder, a cross-modal aligner, and dual decoding heads for simultaneous background audio and speech generation. Specifically, our proposed cross-modal aligner employs causal and non-causal attention mechanisms to improve synchronization and acoustic harmony. Besides, to handle data scarcity, we design a curriculum learning strategy that progressively builds the multimodal capability. Finally, we introduce DualBench, the first benchmark for V2ST evaluation with a carefully curated test set and comprehensive metrics. Experimental results demonstrate that DualDub achieves state-of-the-art performance, generating high-quality and well-synchronized soundtracks with both speech and background audio. DualBench and generated samples of DualDub are available at https://github.com/wjtian-wonderful/DualBench.

Original languageEnglish
Title of host publicationMM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
PublisherAssociation for Computing Machinery, Inc
Pages10671-10680
Number of pages10
ISBN (Electronic)9798400720352
DOIs
StatePublished - 27 Oct 2025
Event33rd ACM International Conference on Multimedia, MM 2025 - Dublin, Ireland
Duration: 27 Oct 202531 Oct 2025

Publication series

NameMM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025

Conference

Conference33rd ACM International Conference on Multimedia, MM 2025
Country/TerritoryIreland
CityDublin
Period27/10/2531/10/25

Keywords

  • curriculum learning
  • data scarcity
  • multimodal alignment
  • soundtrack generation

Fingerprint

Dive into the research topics of 'DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis'. Together they form a unique fingerprint.

Cite this