most citedHow to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition

39 citations · 39 across the 2 of their papers we have counts for

collaborators

5 papers

eess.AS2020

Learning to Count Words in Fluent Speech enables Online Speech Recognition

George Sterpu, Christian Saam, Naomi Harte

Sequence to Sequence models, in particular the Transformer, achieve state of the art results in Automatic Speech Recognition. Practical usage is however limited to cases where full…

eess.AS2020

Should we hard-code the recurrence concept or learn it instead ? Exploring the Transformer architecture for Audio-Visual Speech Recognition

George Sterpu, Christian Saam, Naomi Harte

The audio-visual speech fusion strategy AV Align has shown significant performance improvements in audio-visual speech recognition (AVSR) on the challenging LRS2 dataset. Performan…

eess.AS202039 cited

How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition

George Sterpu, Christian Saam, Naomi Harte

Audio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently propose…

eess.AS2018

Attention-based Audio-Visual Fusion for Robust Automatic Speech Recognition

George Sterpu, Christian Saam, Naomi Harte

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noi…

eess.IV2018

Can DNNs Learn to Lipread Full Sentences?

George Sterpu, Christian Saam, Naomi Harte

Finding visual features and suitable models for lipreading tasks that are more complex than a well-constrained vocabulary has proven challenging. This paper explores state-of-the-a…