97 citations · 308 across the 17 of their papers we have counts for
16 papers · 1 filter
The VoxCeleb Speaker Recognition Challenge: A Retrospective
Jaesung Huh, Joon Son Chung, Arsha Nagrani +4
The VoxCeleb Speaker Recognition Challenges (VoxSRC) were a series of challenges and workshops that ran annually from 2019 to 2023. The challenges primarily evaluated the tasks of…
OxfordVGG Submission to the EGO4D AV Transcription Challenge
Jaesung Huh, Max Bain, Andrew Zisserman
This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX…
VoxSRC 2022: The Fourth VoxCeleb Speaker Recognition Challenge
Jaesung Huh, Andrew Brown, Jee-weon Jung +4
This paper summarises the findings from the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22), which was held in conjunction with INTERSPEECH 2022. The goal of this challenge…
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
Max Bain, Jaesung Huh, Tengda Han +1
Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages. However, their a…
Epic-Sounds: A Large-scale Dataset of Actions That Sound
Jaesung Huh, Jacob Chalk, Evangelos Kazakos +2
We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an ann…
In search of strong embedding extractors for speaker diarisation
Jee-weon Jung, Hee-Soo Heo, Bong-Jin Lee +5
Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several cha…