摘要
Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76% on the AMI test set, demonstrating its robustness and effectiveness in OSD.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 1653-1657 |
| 页数 | 5 |
| 期刊 | Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH |
| DOI | |
| 出版状态 | 已出版 - 2025 |
| 活动 | 26th Interspeech Conference 2025 - Rotterdam, 荷兰 期限: 17 8月 2025 → 21 8月 2025 |
学术指纹
探究 'Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver