activity
20132022
most citedTraining Strategies for Improved Lip-reading

59 citations · 246 across the 25 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS20204 cited

Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision

Abhinav Shukla, Stavros Petridis, Maja Pantic

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual t…

eess.AS2020

Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?

Abhinav Shukla, Stavros Petridis, Maja Pantic

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work…

eess.AS2020

Visually Guided Self Supervised Learning of Speech Representations

Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma +2

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particu…

eess.AS20191 cited

Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition

Pingchuan Ma, Stavros Petridis, Maja Pantic

Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all…

eess.AS2019

Video-Driven Speech Reconstruction using Generative Adversarial Networks

Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis +1

Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. I…

eess.AS2018

End-to-End Speech-Driven Facial Animation with Temporal GANs

Konstantinos Vougioukas, Stavros Petridis, Maja Pantic

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from…