59 citations · 246 across the 25 of their papers we have counts for
6 papers · 1 filter
Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision
Abhinav Shukla, Stavros Petridis, Maja Pantic
The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual t…
Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?
Abhinav Shukla, Stavros Petridis, Maja Pantic
Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work…
Visually Guided Self Supervised Learning of Speech Representations
Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma +2
Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particu…
Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition
Pingchuan Ma, Stavros Petridis, Maja Pantic
Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all…
Video-Driven Speech Reconstruction using Generative Adversarial Networks
Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis +1
Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. I…
End-to-End Speech-Driven Facial Animation with Temporal GANs
Konstantinos Vougioukas, Stavros Petridis, Maja Pantic
Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from…