22 citations · 45 across the 5 of their papers we have counts for
3 papers · 1 filter
Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision
Abhinav Shukla, Stavros Petridis, Maja Pantic
The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual t…
Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?
Abhinav Shukla, Stavros Petridis, Maja Pantic
Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work…
Visually Guided Self Supervised Learning of Speech Representations
Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma +2
Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particu…