VoxCeleb: a large-scale speaker identification dataset
arXiv:1706.08612 · doi:10.21437/Interspeech.2017-950
Abstract
Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent speaker identification dataset collected 'in the wild'. We make two contributions. First, we propose a fully automated pipeline based on computer vision techniques to create the dataset from open-source media. Our pipeline involves obtaining videos from YouTube; performing active speaker verification using a two-stream synchronization Convolutional Neural Network (CNN), and confirming the identity of the speaker using CNN based facial recognition. We use this pipeline to curate VoxCeleb which contains hundreds of thousands of 'real world' utterances for over 1,000 celebrities. Our second contribution is to apply and compare various state of the art speaker identification techniques on our dataset to establish baseline performance. We show that a CNN based architecture obtains the best performance for both identification and verification.
The dataset can be downloaded from http://www.robots.ox.ac.uk/~vgg/data/voxceleb . 1706.08612v2: minor fixes; 6 pages
Cited by in corpus (322)
- VoxCeleb2: Deep Speaker Recognition
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
- Attentive Statistics Pooling for Deep Speaker Embedding
- SpeechBrain: A General-Purpose Speech Toolkit
- Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
- In defence of metric learning for speaker recognition
- Self-Supervised Speech Representation Learning: A Review
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
- ASVspoof 2019: spoofing countermeasures for the detection of synthesized, converted and replayed speech
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- Spot the conversation: speaker diarisation in the wild
- Towards Learning a Universal Non-Semantic Representation of Speech
- Neural Head Reenactment with Latent Pose Descriptors
- Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals
- Fine-tuning wav2vec2 for speaker recognition
- ECAPA-TDNN Embeddings for Speaker Diarization
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- Interpretable Convolutional Filters with SincNet
- Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling
- Emotional Speech-Driven Animation with Content-Emotion Disentanglement
- Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information
- X-vectors: New Quantitative Biomarkers for Early Parkinson's Disease Detection from Speech
- A Good Image Generator Is What You Need for High-Resolution Video Synthesis
- VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
- Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification
- BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations
- The IDLAB VoxSRC-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker Verification
- Meta-TTS: Meta-Learning for Few-Shot Speaker Adaptive Text-to-Speech
- Neural Predictive Coding using Convolutional Neural Networks towards Unsupervised Learning of Speaker Characteristics
- WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?
- VoxSRC 2019: The first VoxCeleb Speaker Recognition Challenge
- Optimization of data-driven filterbank for automatic speaker verification
- Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances
- Deep Learning in Physical Layer: Review on Data Driven End-to-End Communication Systems and their Enabling Semantic Applications
- TRILLsson: Distilled Universal Paralinguistic Speech Representations
- Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision
- Talking Face Generation by Adversarially Disentangled Audio-Visual Representation
- The SpeakIn System for VoxCeleb Speaker Recognition Challange 2021
- Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis
- Voice activity detection in the wild: A data-driven approach using teacher-student training
- Everybody's Talkin': Let Me Talk as You Want
- Universal Paralinguistic Speech Representations Using Self-Supervised Conformers
- Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification
- Deep learning methods in speaker recognition: a review
- AVA-AVD: Audio-Visual Speaker Diarization in the Wild
- Deep multi-metric learning for text-independent speaker verification
- SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
- AM-MobileNet1D: A Portable Model for Speaker Recognition
- SkipConvNet: Skip Convolutional Neural Network for Speech Dereverberation using Optimally Smoothed Spectral Mapping
- Semi-Supervised Contrastive Learning with Generalized Contrastive Loss and Its Application to Speaker Recognition
- Improving Zero-shot Voice Style Transfer via Disentangled Representation Learning
- Hierarchical Cross-Modal Talking Face Generationwith Dynamic Pixel-Wise Loss
- Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
- Scalable and Efficient Neural Speech Coding: A Hybrid Design
- The VoicePrivacy 2020 Challenge Evaluation Plan
- ATST: Audio Representation Learning with Teacher-Student Transformer
- On the Resilience of Biometric Authentication Systems against Random Inputs
- iQIYI-VID: A Large Dataset for Multi-modal Person Identification
- Attention Back-end for Automatic Speaker Verification with Multiple Enrollment Utterances
- Families In Wild Multimedia: A Multimodal Database for Recognizing Kinship
- NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
- The DKU-DukeECE Systems for VoxCeleb Speaker Recognition Challenge 2020
- A Deep Network for Arousal-Valence Emotion Prediction with Acoustic-Visual Cues
- The VoxCeleb Speaker Recognition Challenge: A Retrospective
- Defend Data Poisoning Attacks on Voice Authentication
- Multimodal Self-Supervised Learning of General Audio Representations
- Talking Face Generation by Conditional Recurrent Adversarial Network
- USTC-NELSLIP System Description for DIHARD-III Challenge
- The Influence of Dataset Partitioning on Dysfluency Detection Systems
- A Unified Deep Learning Framework for Short-Duration Speaker Verification in Adverse Environments
- Graph-based Label Propagation for Semi-Supervised Speaker Identification
- Cross-Lingual Speaker Verification with Domain-Balanced Hard Prototype Mining and Language-Dependent Score Normalization
- AVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection
- Non-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech
- Combination of Deep Speaker Embeddings for Diarisation
- FRILL: A Non-Semantic Speech Embedding for Mobile Devices
- FreeStyleGAN: Free-view Editable Portrait Rendering with the Camera Manifold
- Self-supervised Representation Learning With Path Integral Clustering For Speaker Diarization
- Advances in Online Audio-Visual Meeting Transcription
- Channel adversarial training for speaker verification and diarization
- Utterance-level Aggregation For Speaker Recognition In The Wild
- A Real-time Speaker Diarization System Based on Spatial Spectrum
- Deep Audio-Visual Learning: A Survey
- Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
- Am I a Real or Fake Celebrity? Measuring Commercial Face Recognition Web APIs under Deepfake Impersonation Attack
- Meeting Transcription Using Virtual Microphone Arrays
- DECAR: Deep Clustering for learning general-purpose Audio Representations
- DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
- Speaker Anonymization with Distribution-Preserving X-Vector Generation for the VoicePrivacy Challenge 2020
- The effect of speech pathology on automatic speaker verification -- a large-scale study
- Disentangled speaker and nuisance attribute embedding for robust speaker verification
- Smartphone Multi-modal Biometric Authentication: Database and Evaluation
- Emotion Recognition in Speech using Cross-Modal Transfer in the Wild
- Lightweight Speaker Verification for Online Identification of New Speakers with Short Segments
- Exploring wav2vec 2.0 on speaker verification and language identification
- AutoSpeech: Neural Architecture Search for Speaker Recognition
- Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker Identification
- From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
- Improving fairness in speaker verification via Group-adapted Fusion Network
- CN-CELEB: a challenging Chinese speaker recognition dataset
- ResNeXt and Res2Net Structures for Speaker Verification
- Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
- Large-scale multilingual audio visual dubbing
- HarperValleyBank: A Domain-Specific Spoken Dialog Corpus
- Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
- Adversarial Attacks on GMM i-vector based Speaker Verification Systems
- Learning spectro-temporal representations of complex sounds with parameterized neural networks
- openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
- Speaker Diarization as a Fully Online Learning Problem in MiniVox
- Emotional Prosody Control for Speech Generation
- VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency
- Conditional Injective Flows for Bayesian Imaging
- The DKU-DukeECE System for the Self-Supervision Speaker Verification Task of the 2021 VoxCeleb Speaker Recognition Challenge
- Joint gender and age estimation based on speech signals using x-vectors and transfer learning
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering
- Generalizing Speaker Verification for Spoof Awareness in the Embedding Space
- Coding Speech through Vocal Tract Kinematics
- Fake the Real: Backdoor Attack on Deep Speech Classification via Voice Conversion
- ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
- Contrastive-mixup learning for improved speaker verification
- An Adaptive X-vector Model for Text-independent Speaker Verification
- Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
- Deep Normalization for Speaker Vectors
- FluentNet: End-to-End Detection of Speech Disfluency with Deep Learning
- Adversarial defense for automatic speaker verification by cascaded self-supervised learning models
- Wav2Pix: Speech-conditioned Face Generation using Generative Adversarial Networks
- Training speaker recognition systems with limited data
- Personal VAD: Speaker-Conditioned Voice Activity Detection
- MAAS: Multi-modal Assignation for Active Speaker Detection
- CN-Celeb: multi-genre speaker recognition
- Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
- Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings
- DIHARD II is Still Hard: Experimental Results and Discussions from the DKU-LENOVO Team
- A Comparison of Metric Learning Loss Functions for End-To-End Speaker Verification
- Fully Supervised Speaker Diarization
- Robust Speaker Recognition Using Speech Enhancement And Attention Model
- Complex-Valued Time-Frequency Self-Attention for Speech Dereverberation
- Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
- What all do audio transformer models hear? Probing Acoustic Representations for Language Delivery and its Structure
- SLMIA-SR: Speaker-Level Membership Inference Attacks against Speaker Recognition Systems
- Spectrogram features for audio and speech analysis
- TSUP Speaker Diarization System for Conversational Short-phrase Speaker Diarization Challenge
- Semi-supervised acoustic modelling for five-lingual code-switched ASR using automatically-segmented soap opera speech
- Cosine Scoring with Uncertainty for Neural Speaker Embedding
- Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
- T: Multi-Modal Continuous Valence-Arousal Estimation in the Wild
- T-vectors: Weakly Supervised Speaker Identification Using Hierarchical Transformer Model
- Bayesian x-vector: Bayesian Neural Network based x-vector System for Speaker Verification
- Voice-Indistinguishability: Protecting Voiceprint in Privacy-Preserving Speech Data Release
- VGGSound: A Large-scale Audio-Visual Dataset
- Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information
- Voice Cloning: a Multi-Speaker Text-to-Speech Synthesis Approach based on Transfer Learning
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording
- LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification
- Versatile Face Animator: Driving Arbitrary 3D Facial Avatar in RGBD Space
- AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
- TwoStreamVAN: Improving Motion Modeling in Video Generation
- Privacy-Protecting Techniques for Behavioral Biometric Data: A Survey
- Learning Audio-Visual Dereverberation
- Federated Learning of User Authentication Models
- Defending Your Voice: Adversarial Attack on Voice Conversion
- Understanding the Tradeoffs in Client-side Privacy for Downstream Speech Tasks
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- Speaker De-identification System using Autoencoders and Adversarial Training
- Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
- Diffusion-Based Adversarial Purification for Speaker Verification
- Adversarial Reweighting for Speaker Verification Fairness
- A Multi-View Approach To Audio-Visual Speaker Verification
- Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification
- Bio-Inspired Modality Fusion for Active Speaker Detection
- Deep Speaker Embeddings for Far-Field Speaker Recognition on Short Utterances
- An Investigation of the Effectiveness of Phase for Audio Classification
- Improving Speaker Identification for Shared Devices by Adapting Embeddings to Speaker Subsets
- Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation
- Optimizing Multi-Taper Features for Deep Speaker Verification
- A Study of Few-Shot Audio Classification
- AniFaceDiff: Animating Stylized Avatars via Parametric Conditioned Diffusion Models
- N-HANS: Introducing the Augsburg Neuro-Holistic Audio-eNhancement System
- VAE-based regularization for deep speaker embedding
- Scalable Data Annotation Pipeline for High-Quality Large Speech Datasets Development
- Pairwise Discriminative Neural PLDA for Speaker Verification
- Deep Speaker Vector Normalization with Maximum Gaussianality Training
- End-To-End Audiovisual Feature Fusion for Active Speaker Detection
- PriorityCut: Occlusion-guided Regularization for Warp-based Image Animation
- Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
- Voting for the right answer: Adversarial defense for speaker verification
- WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
- Online Speaker Diarization with Relation Network
- A Memory Augmented Architecture for Continuous Speaker Identification in Meetings
- Experimenting with Additive Margins for Contrastive Self-Supervised Speaker Verification
- Understanding Semantics from Speech Through Pre-training
- An iterative framework for self-supervised deep speaker representation learning
- Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
- Label-efficient audio classification through multitask learning and self-supervision
- LEAP System for SRE19 CTS Challenge -- Improvements and Error Analysis
- Improved Meta-Learning Training for Speaker Verification
- Investigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordings
- North America Bixby Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021
- Learnable MFCCs for Speaker Verification
- Intra-class variation reduction of speaker representation in disentanglement framework
- Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
- GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
- Generative Human Video Compression with Multi-granularity Temporal Trajectory Factorization
- Single Channel Far Field Feature Enhancement For Speaker Verification In The Wild
- Towards Learning Universal Audio Representations
- asya: Mindful verbal communication using deep learning
- Federated Learning of User Verification Models Without Sharing Embeddings
- SEEF-ALDR: A Speaker Embedding Enhancement Framework via Adversarial Learning based Disentangled Representation
- Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR
- The HCCL Speaker Verification System for Far-Field Speaker Verification Challenge
- Learning to Aggregate and Personalize 3D Face from In-the-Wild Photo Collection
- Significance of Speaker Embeddings and Temporal Context for Depression Detection
- Self-supervised learning for audio-visual speaker diarization
- SpeechNAS: Towards Better Trade-off between Latency and Accuracy for Large-Scale Speaker Verification
- A Unified Deep Speaker Embedding Framework for Mixed-Bandwidth Speech Data
- Towards Measuring and Scoring Speaker Diarization Fairness
- Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
- HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
- Structural sparsification for Far-field Speaker Recognition with GNA
- Text-Independent Speaker Verification with Dual Attention Network
- Speaker Diarization with Region Proposal Network
- ET-GAN: Cross-Language Emotion Transfer Based on Cycle-Consistent Generative Adversarial Networks
- Rep Works in Speaker Verification
- A Study on Angular Based Embedding Learning for Text-independent Speaker Verification
- The sound of my voice: speaker representation loss for target voice separation
- FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation Model
- An end-to-end approach for the verification problem: learning the right distance
- FaceEraser: Removing Facial Parts for Augmented Reality
- Combining Attention with Flow for Person Image Synthesis
- Multi-modal Residual Perceptron Network for Audio-Video Emotion Recognition
- Practical Attacks on Voice Spoofing Countermeasures
- A Tandem Framework Balancing Privacy and Security for Voice User Interfaces
- Serialized Multi-Layer Multi-Head Attention for Neural Speaker Embedding
- MACCIF-TDNN: Multi aspect aggregation of channel and context interdependence features in TDNN-based speaker verification
- A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio
- QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus
- Adaptive Margin Circle Loss for Speaker Verification
- Robust Acoustic Domain Identification with its Application to Speaker Diarization
- Speaker Characterization by means of Attention Pooling
- Multi-Task Learning with High-Order Statistics for X-vector based Text-Independent Speaker Verification
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
- An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
- Self-Tuning Spectral Clustering for Speaker Diarization
- Deep generative LDA
- DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning
- Graph-based Multi-View Fusion and Local Adaptation: Mitigating Within-Household Confusability for Speaker Identification
- Automatic Face Understanding: Recognizing Families in Photos
- The DKU-Duke-Lenovo System Description for the Third DIHARD Speech Diarization Challenge
- The DKU System for the Speaker Recognition Task of the 2019 VOiCES from a Distance Challenge
- A Study on Decoupled Probabilistic Linear Discriminant Analysis
- Diarization of Legal Proceedings. Identifying and Transcribing Judicial Speech from Recorded Court Audio
- Temporal Dynamic Convolutional Neural Network for Text-Independent Speaker Verification and Phonemetic Analysis
- Speaker Re-identification with Speaker Dependent Speech Enhancement
- SVEva Fair: A Framework for Evaluating Fairness in Speaker Verification
- Weakly Supervised Training of Hierarchical Attention Networks for Speaker Identification
- LSTM and GPT-2 Synthetic Speech Transfer Learning for Speaker Recognition to Overcome Data Scarcity
- The HUAWEI Speaker Diarisation System for the VoxCeleb Speaker Diarisation Challenge
- Unsupervised Representation Learning for Speaker Recognition via Contrastive Equilibrium Learning
- Momentum Contrast Speaker Representation Learning
- HLT-NUS Submission for NIST 2019 Multimedia Speaker Recognition Evaluation
- Compositional embedding models for speaker identification and diarization with simultaneous speech from 2+ speakers
- Open-set Short Utterance Forensic Speaker Verification using Teacher-Student Network with Explicit Inductive Bias
- Delving into VoxCeleb: environment invariant speaker recognition
- Mixture factorized auto-encoder for unsupervised hierarchical deep factorization of speech signal
- Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification
- Cross-modal Speaker Verification and Recognition: A Multilingual Perspective
- Learning Inverse Rendering of Faces from Real-world Videos
- HI-MIA : A Far-field Text-Dependent Speaker Verification Database and the Baselines
- The DKU-SMIIP System for NIST 2018 Speaker Recognition Evaluation
- DropClass and DropAdapt: Dropping classes for deep speaker representation learning
- Gaussian-Constrained training for speaker verification
- Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study
- Poformer: A simple pooling transformer for speaker verification
- Speaker Recognition with Random Digit Strings Using Uncertainty Normalized HMM-based i-vectors
- Few Shot Text-Independent speaker verification using 3D-CNN
- Online Speaker Adaptation for WaveNet-based Neural Vocoders
- Voice Biometrics Security: Extrapolating False Alarm Rate via Hierarchical Bayesian Modeling of Speaker Verification Scores
- The Impact of Spatiotemporal Augmentations on Self-Supervised Audiovisual Representation Learning
- An empirical study of weakly supervised audio tagging embeddings for general audio representations
- Active Speakers in Context
- Learning Metrics from Mean Teacher: A Supervised Learning Method for Improving the Generalization of Speaker Verification System
- Streaming Multi-talker Speech Recognition with Joint Speaker Identification
- Large-scale Multi-modal Person Identification in Real Unconstrained Environments
- Inclusive Speaker Verification with Adaptive thresholding
- Gaussian speaker embedding learning for text-independent speaker verification
- Theoretical Guarantees of Deep Embedding Losses Under Label Noise
- An MAP Estimation for Between-Class Variance
- Self-Supervised 3D Face Reconstruction via Conditional Estimation
- Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph
- Duality Temporal-channel-frequency Attention Enhanced Speaker Representation Learning
- One-shot Face Reenactment Using Appearance Adaptive Normalization
- Supervised attention for speaker recognition
- Sparse Architectures for Text-Independent Speaker Verification Using Deep Neural Networks
- An initial investigation on optimizing tandem speaker verification and countermeasure systems using reinforcement learning
- Exploring End-to-End Multi-channel ASR with Bias Information for Meeting Transcription
- Exploring Voice Conversion based Data Augmentation in Text-Dependent Speaker Verification
- MCSAE: Masked Cross Self-Attentive Encoding for Speaker Embedding
- Comparison of Speaker Role Recognition and Speaker Enrollment Protocol for conversational Clinical Interviews
- Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling
- An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
- VAE-based Domain Adaptation for Speaker Verification
- Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification
- Collaborative Feature Aggregation for Face Super-Resolution and Robust Re-Identification
- Within-sample variability-invariant loss for robust speaker recognition under noisy environments
- SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
- When Automatic Voice Disguise Meets Automatic Speaker Verification
- Multi-task Metric Learning for Text-independent Speaker Verification
- Squeezing value of cross-domain labels: a decoupled scoring approach for speaker verification
- The Newsbridge -Telecom SudParis VoxCeleb Speaker Recognition Challenge 2022 System Description
- Utterance Clustering Using Stereo Audio Channels
- Tongji University Undergraduate Team for the VoxCeleb Speaker Recognition Challenge2020
- Cross-domain Adaptation with Discrepancy Minimization for Text-independent Forensic Speaker Verification
- Estimating Uniqueness of I-Vector Representation of Human Voice
- Non-native Speaker Verification for Spoken Language Assessment
- How Far Are We from Robust Voice Conversion: A Survey
- Cross-lingual Transfer for Speech Processing using Acoustic Language Similarity
- Non-local convolutional neural networks (nlcnn) for speaker recognition
- Speaker Representation Learning using Global Context Guided Channel and Time-Frequency Transformations
- Sparse to Dense Motion Transfer for Face Image Animation