Lip Reading Sentences in the Wild
arXiv:1611.05358 · doi:10.1109/CVPR.2017.367
Abstract
The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an open-world problem - unconstrained natural language sentences, and in the wild videos. Our key contributions are: (1) a 'Watch, Listen, Attend and Spell' (WLAS) network that learns to transcribe videos of mouth motion to characters; (2) a curriculum learning strategy to accelerate training and to reduce overfitting; (3) a 'Lip Reading Sentences' (LRS) dataset for visual speech recognition, consisting of over 100,000 natural sentences from British television. The WLAS model trained on the LRS dataset surpasses the performance of all previous work on standard lip reading benchmark datasets, often by a significant margin. This lip reading performance beats a professional lip reader on videos from BBC television, and we also demonstrate that visual information helps to improve speech recognition performance even when the audio is available.
Cited by in corpus (27)
- Deep Audio-Visual Speech Recognition
- Visual Speech Recognition for Multiple Languages in the Wild
- Self-Supervised Learning for Videos: A Survey
- Silent Speech Interfaces for Speech Restoration: A Review
- Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
- Multi-modal Multi-channel Target Speech Separation
- Neural Head Reenactment with Latent Pose Descriptors
- Perfect match: Improved cross-modal embeddings for audio-visual synchronisation
- Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
- Lip-reading with Densely Connected Temporal Convolutional Networks
- Temporally Guided Music-to-Body-Movement Generation
- How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition
- LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading
- Context-self contrastive pretraining for crop type semantic segmentation
- Families In Wild Multimedia: A Multimodal Database for Recognizing Kinship
- DualLip: A System for Joint Lip Reading and Generation
- Scaling up sign spotting through sign language dictionaries
- SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
- Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
- Automated Speaker Independent Visual Speech Recognition: A Comprehensive Survey
- Towards Accurate Lip-to-Speech Synthesis in-the-Wild
- Uncovering the Visual Contribution in Audio-Visual Speech Recognition
- FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
- Word-level Persian Lipreading Dataset
- TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
- Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions