12 citations · 16 across the 14 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023
Joint Audio and Speech Understanding
Yuan Gong, Alexander H. Liu, Hongyin Luo +2
Humans are surrounded by audio signals that include both speech and non-speech sounds. The recognition and understanding of speech and non-speech audio events, along with a profoun…
cs.SD2023
Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers
Yuan Gong, Sameer Khurana, Leonid Karlinsky +1
In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show…