7 papers
openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
Kishan K C, Zhenning Tan, Long Chen +4
Household speaker identification with few enrollment utterances is an important yet challenging problem, especially when household members share similar voice characteristics and r…
Improving fairness in speaker verification via Group-adapted Fusion Network
Hua Shen, Yuguang Yang, Guoli Sun +4
Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typical…
Contrastive-mixup learning for improved speaker verification
Xin Zhang, Minho Jin, Roger Cheng +3
This paper proposes a novel formulation of prototypical loss with mixup for speaker verification. Mixup is a simple yet efficient data augmentation technique that fabricates a weig…
ASR-Aware End-to-end Neural Diarization
Aparna Khare, Eunjung Han, Yuguang Yang +1
We present a Conformer-based end-to-end neural diarization (EEND) model that uses both acoustic input and features derived from an automatic speech recognition (ASR) model. Two cat…
Improving Speaker Identification for Shared Devices by Adapting Embeddings to Speaker Subsets
Zhenning Tan, Yuguang Yang, Eunjung Han +1
Speaker identification typically involves three stages. First, a front-end speaker embedding model is trained to embed utterance and speaker profiles. Second, a scoring function is…
End-to-end Neural Diarization: From Transformer to Conformer
Yi Chieh Liu, Eunjung Han, Chul Lee +1
We propose a new end-to-end neural diarization (EEND) system that is based on Conformer, a recently proposed neural architecture that combines convolutional mappings and Transforme…