97 citations · 280 across the 17 of their papers we have counts for
15 papers · 1 filter
FlexiAST: Flexibility is What AST Needs
Jiu Feng, Mehmet Hamza Erol, Joon Son Chung +1
The objective of this work is to give patch-size flexibility to Audio Spectrogram Transformers (AST). Recent advancements in ASTs have shown superior performance in various audio-b…
VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
Arsha Nagrani, Joon Son Chung, Jaesung Huh +6
We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker…
Supervised attention for speaker recognition
Seong Min Kye, Joon Son Chung, Hoirin Kim
The recently proposed self-attentive pooling (SAP) has shown good performance in several speaker recognition systems. In SAP systems, the context vector is trained end-to-end toget…
Look who's not talking
Youngki Kwon, Hee Soo Heo, Jaesung Huh +2
The objective of this work is speaker diarisation of speech recordings 'in the wild'. The ability to determine speech segments is a crucial part of diarisation systems, accounting…
The ins and outs of speaker recognition: lessons from VoxSRC 2020
Yoohwan Kwon, Hee-Soo Heo, Bong-Jin Lee +1
The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing differen…
Playing a Part: Speaker Verification at the Movies
Andrew Brown, Jaesung Huh, Arsha Nagrani +2
The goal of this work is to investigate the performance of popular speaker recognition models on speech segments from movies, where often actors intentionally disguise their voice…