跳到主要导航 跳到搜索 跳到主要内容

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

  • Ming Gao
  • , Shilong Wu
  • , Hang Chen
  • , Jun Du
  • , Chin Hui Lee
  • , Shinji Watanabe
  • , Jingdong Chen
  • , Siniscalchi Sabato Marco
  • , Odette Scharenborg
  • University of Science and Technology of China
  • Georgia Institute of Technology
  • Carnegie Mellon University
  • University of Palermo
  • Delft University of Technology

科研成果: 期刊稿件会议文章同行评审

摘要

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.

源语言英语
页(从-至)1888-1892
页数5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
DOI
出版状态已出版 - 2025
活动26th Interspeech Conference 2025 - Rotterdam, 荷兰
期限: 17 8月 202521 8月 2025

指纹

探究 'The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition' 的科研主题。它们共同构成独一无二的指纹。

引用此