97 citations · 246 across the 8 of their papers we have counts for
13 papers
In search of strong embedding extractors for speaker diarisation
Jee-weon Jung, Hee-Soo Heo, Bong-Jin Lee +5
Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several cha…
With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition
Evangelos Kazakos, Jaesung Huh, Arsha Nagrani +2
In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order t…
VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
Arsha Nagrani, Joon Son Chung, Jaesung Huh +6
We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker…
Look who's not talking
Youngki Kwon, Hee Soo Heo, Jaesung Huh +2
The objective of this work is speaker diarisation of speech recordings 'in the wild'. The ability to determine speech segments is a crucial part of diarisation systems, accounting…
Playing a Part: Speaker Verification at the Movies
Andrew Brown, Jaesung Huh, Arsha Nagrani +2
The goal of this work is to investigate the performance of popular speaker recognition models on speech segments from movies, where often actors intentionally disguise their voice…
Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
Hee Soo Heo, Bong-Jin Lee, Jaesung Huh +1
This report describes our submission to the VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020. We perform a careful analysis of speaker recognition models based o…