31 citations · 58 across the 4 of their papers we have counts for
8 papers
Interactive decoding of words from visual speech recognition models
Brendan Shillingford, Yannis Assael, Misha Denil
This work describes an interactive decoding method to improve the performance of visual speech recognition systems using user input to compensate for the inherent ambiguity of the…
Large-scale multilingual audio visual dubbing
Yi Yang, Brendan Shillingford, Yannis Assael +9
We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed…
Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
Takaki Makino, Hank Liao, Yannis Assael +4
This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a sy…
Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini +2
To increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for the…
Speech bandwidth extension with WaveNet
Archit Gupta, Brendan Shillingford, Yannis Assael +1
Large-scale mobile communication systems tend to contain legacy transmission channels with narrowband bottlenecks, resulting in characteristic "telephone-quality" audio. While high…
Sample Efficient Adaptive Text-to-Speech
Yutian Chen, Yannis Assael, Brendan Shillingford +11
We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and ind…