BUT System Description to VoxCeleb Speaker Recognition Challenge 2019
arXiv:1910.12592
Abstract
In this report, we describe the submission of Brno University of Technology (BUT) team to the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2019. We also provide a brief analysis of different systems on VoxCeleb-1 test sets. Submitted systems for both Fixed and Open conditions are a fusion of 4 Convolutional Neural Network (CNN) topologies. The first and second networks have ResNet34 topology and use two-dimensional CNNs. The last two networks are one-dimensional CNN and are based on the x-vector extraction topology. Some of the networks are fine-tuned using additive margin angular softmax. Kaldi FBanks and Kaldi PLPs were used as features. The difference between Fixed and Open systems lies in the used training data and fusion strategy. The best systems for Fixed and Open conditions achieved 1.42% and 1.26% ERR on the challenge evaluation set respectively.
Cited by in corpus (21)
- SpeechBrain: A General-Purpose Speech Toolkit
- Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
- VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
- VoxSRC 2019: The first VoxCeleb Speaker Recognition Challenge
- Self-Supervised Learning Based Domain Adaptation for Robust Speaker Verification
- The SpeakIn System for VoxCeleb Speaker Recognition Challange 2021
- Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
- Adversarial defense for automatic speaker verification by cascaded self-supervised learning models
- The xx205 System for the VoxCeleb Speaker Recognition Challenge 2020
- A Multi-View Approach To Audio-Visual Speaker Verification
- Deep Speaker Embeddings for Far-Field Speaker Recognition on Short Utterances
- Domain-Invariant Speaker Vector Projection by Model-Agnostic Meta-Learning
- Rep Works in Speaker Verification
- Delving into VoxCeleb: environment invariant speaker recognition
- LSTM and GPT-2 Synthetic Speech Transfer Learning for Speaker Recognition to Overcome Data Scarcity
- MACCIF-TDNN: Multi aspect aggregation of channel and context interdependence features in TDNN-based speaker verification
- Duality Temporal-channel-frequency Attention Enhanced Speaker Representation Learning
- STC speaker recognition systems for the NIST SRE 2021
- Xi-Vector Embedding for Speaker Recognition
- Towards Robust Speaker Verification with Target Speaker Enhancement
- Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph