Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
arXiv:2009.14153
Abstract
This report describes our submission to the VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020. We perform a careful analysis of speaker recognition models based on the popular ResNet architecture, and train a number of variants using a range of loss functions. Our results show significant improvements over most existing works without the use of model ensemble or post-processing. We release the training code and pre-trained models as unofficial baselines for this year's challenge.
References in corpus (4)
Cited by in corpus (5)
- Sparsely Overlapped Speech Training in the Time Domain: Joint Learning of Target Speech Separation and Personal VAD Benefits
- ShaneRun System Description to VoxCeleb Speaker Recognition Challenge 2020
- Generalized Operating Procedure for Deep Learning: an Unconstrained Optimal Design Perspective
- Duality Temporal-channel-frequency Attention Enhanced Speaker Representation Learning
- Multi-Level Transfer Learning from Near-Field to Far-Field Speaker Verification