3 papers
eess.AS2024
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Andrew Rouditchenko, Yuan Gong, Samuel Thomas +4
Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is…
eess.AS2024
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
Alexander H. Liu, Sung-Lin Yeh, James Glass
Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. Howev…
cs.SD2023
Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers
Yuan Gong, Sameer Khurana, Leonid Karlinsky +1
In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show…