4 papers
Epic-Sounds: A Large-scale Dataset of Actions That Sound
Jaesung Huh, Jacob Chalk, Evangelos Kazakos +2
We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an ann…
Character-aware audio-visual subtitling in context
Jaesung Huh, Andrew Zisserman
This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character re…
The VoxCeleb Speaker Recognition Challenge: A Retrospective
Jaesung Huh, Joon Son Chung, Arsha Nagrani +4
The VoxCeleb Speaker Recognition Challenges (VoxSRC) were a series of challenges and workshops that ran annually from 2019 to 2023. The challenges primarily evaluated the tasks of…
TIM: A Time Interval Machine for Audio-Visual Action Recognition
Jacob Chalk, Jaesung Huh, Evangelos Kazakos +2
Diverse actions give rise to rich audio-visual signals in long videos. Recent works showcase that the two modalities of audio and video exhibit different temporal extents of events…