61 citations · 80 across the 7 of their papers we have counts for
3 papers · 1 filter
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Andrew Rouditchenko, Yuan Gong, Samuel Thomas +4
Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is…
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
Alexander H. Liu, Sung-Lin Yeh, James Glass
Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. Howev…
Listen, Think, and Understand
Yuan Gong, Hongyin Luo, Alexander H. Liu +2
The ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications. Although significant progress has been made in this a…