LipNet: End-to-End Sentence-level Lipreading
arXiv:1611.01599
Abstract
Lipreading is the task of decoding text from the movement of a speaker's mouth. Traditional approaches separated the problem into two stages: designing or learning visual features, and prediction. More recent deep lipreading approaches are end-to-end trainable (Wand et al., 2016; Chung & Zisserman, 2016a). However, existing work on models trained end-to-end perform only word classification, rather than sentence-level sequence prediction. Studies have shown that human lipreading performance increases for longer words (Easton & Basala, 1982), indicating the importance of features capturing temporal context in an ambiguous communication channel. Motivated by this observation, we present LipNet, a model that maps a variable-length sequence of video frames to text, making use of spatiotemporal convolutions, a recurrent network, and the connectionist temporal classification loss, trained entirely end-to-end. To the best of our knowledge, LipNet is the first end-to-end sentence-level lipreading model that simultaneously learns spatiotemporal visual features and a sequence model. On the GRID corpus, LipNet achieves 95.2% accuracy in sentence-level, overlapped speaker split task, outperforming experienced human lipreaders and the previous 86.4% word-level state-of-the-art accuracy (Gergen et al., 2016).
References in corpus (2)
Cited by in corpus (35)
- Learn an Effective Lip Reading Model without Pains
- Learning Spatio-Temporal Features with Two-Stream Deep 3D CNNs for Lipreading
- Cortical microcircuits as gated-recurrent neural networks
- Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)
- Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
- LipReading with 3D-2D-CNN BLSTM-HMM and word-CTC models
- Attentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition
- Vid2speech: Speech Reconstruction from Silent Video
- Simulation Experiment of BCI Based on Imagined Speech EEG Decoding
- Visual Keyword Spotting with Attention
- Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- Synchronous Bidirectional Learning for Multilingual Lip Reading
- End-to-end Audio-visual Speech Recognition with Conformers
- Exploring Deep Learning for Joint Audio-Visual Lip Biometrics
- Lip-reading with Hierarchical Pyramidal Convolution and Self-Attention
- Discriminative Multi-modality Speech Recognition
- Advances and Challenges in Deep Lip Reading
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive Memory
- "Notic My Speech" -- Blending Speech Patterns With Multimedia
- Disentangling Homophemes in Lip Reading using Perplexity Analysis
- MobiVSR: A Visual Speech Recognition Solution for Mobile Devices
- Listen, Look and Deliberate: Visual context-aware speech recognition using pre-trained text-video representations
- Towards Pose-invariant Lip-Reading
- Audio-Visual Decision Fusion for WFST-based and seq2seq Models
- Sensor Transformation Attention Networks
- Latent Variable Algorithms for Multimodal Learning and Sensor Fusion
- Continuous Speech Recognition using EEG and Video
- Interactive decoding of words from visual speech recognition models
- One Shot Audio to Animated Video Generation
- An Empirical Analysis of Deep Audio-Visual Models for Speech Recognition
- Robust One Shot Audio to Video Generation
- Multi Modal Adaptive Normalization for Audio to Video Generation
- End-to-end Silent Speech Recognition with Acoustic Sensing
- Predicting Video features from EEG and Vice versa