Beijing ZKJ-NPU Speaker Verification System for VoxCeleb Speaker Recognition Challenge 2021
arXiv:2109.03568
Abstract
In this report, we describe the Beijing ZKJ-NPU team submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We participated in the fully supervised speaker verification track 1 and track 2. In the challenge, we explored various kinds of advanced neural network structures with different pooling layers and objective loss functions. In addition, we introduced the ResNet-DTCF, CoAtNet and PyConv networks to advance the performance of CNN-based speaker embedding model. Moreover, we applied embedding normalization and score normalization at the evaluation stage. By fusing 11 and 14 systems, our final best performances (minDCF/EER) on the evaluation trails are 0.1205/2.8160% and 0.1175/2.8400% respectively for track 1 and 2. With our submission, we came to the second place in the challenge for both tracks.
References in corpus (8)
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- The BOSARIS Toolkit: Theory, Algorithms and Code for Surviving the New DCF
- Pyramidal Convolution: Rethinking Convolutional Neural Networks for Visual Recognition
- Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
- VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
- Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification
- The IDLAB VoxSRC-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker Verification
- VoxSRC 2019: The first VoxCeleb Speaker Recognition Challenge