跳到主要导航 跳到搜索 跳到主要内容

Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder

  • He Wang
  • , Pengcheng Guo
  • , Xucheng Wan
  • , Huan Zhou
  • , Lei Xie
  • Northwestern Polytechnical University Xian
  • Huawei Technologies Co., Ltd.

科研成果: 书/报告/会议事项章节会议稿件同行评审

4 引用 (Scopus)

摘要

Automatic lip-reading (ALR) aims to automatically tran-scribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single scale. In this paper, we propose to enhance lip-reading by incorporating multi-scale video data and multi-encoder. Specifically, we first introduce a novel multi-scale lip motion extraction algorithm based on the size of the speaker's face and propose an Enhanced ResNet3D visual front-end (VFE) to extract lip features at different scales. For the multi-encoder, in addition to the mainstream Transformer and Conformer, we also incorporate the recently proposed Branch-former and E-Branchformer as visual encoders. In the experiments, we explore the influence of different video data scales and encoders on ALR system performance and fuse the texts transcribed by all ALR systems using recognizer output voting error reduction (ROVER). Finally, our proposed approach placed second in the ICME 2024 ChatCLR Challenge Task 2, with a 21.52% reduction in character error rate (CER) compared to the official baseline on the evaluation set.

源语言英语
主期刊名2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798350379815
DOI
出版状态已出版 - 2024
活动2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024 - Niagara Falls, 加拿大
期限: 15 7月 202419 7月 2024

丛书

姓名2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024

会议

会议2024 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2024
国家/地区加拿大
Niagara Falls
时期15/07/2419/07/24

学术指纹

探究 'Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder' 的科研主题。它们共同构成独一无二的学术指纹。

引用此