摘要
This paper presents a robust visual feature based on Visemic LDA for audio visual speech recognition, which captures dynamic lip contour information and reflects the viseme classes of visual speech. The paper also introduces an automatic labeling method using the speech recognition results for LDA training data, which avoids the tedious manually labeling work and labeling errors. Experimental results show that the audio visual speech recognition system based on the visual features presented in this paper can greatly increase the speech recognition rate in noisy conditions. The combination of the visual feature with multi-stream HMM can bring the recognition rate of over 80% at a 10 dB SNR noisy condition.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 64-68 |
| 页数 | 5 |
| 期刊 | Dianzi Yu Xinxi Xuebao/Journal of Electronics and Information Technology |
| 卷 | 27 |
| 期 | 1 |
| 出版状态 | 已出版 - 1月 2005 |
学术指纹
探究 'A robust dynamic mouth feature based on Visemic LDA for audio visual speech recognition' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver