activity
20172021
most citedTime-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker Verification

32 citations · 58 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS20191 cited

MCE 2018: The 1st Multi-target Speaker Detection and Identification Challenge Evaluation

Suwon Shon, Najim Dehak, Douglas Reynolds +1

The Multi-target Challenge aims to assess how well current speech technology is able to determine whether or not a recorded utterance was spoken by one of a large number of blackli…

eess.AS2019

VoiceID Loss: Speech Enhancement for Speaker Verification

Suwon Shon, Hao Tang, James Glass

In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly…

eess.AS2018

Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain

Suwon Shon, Ahmed Ali, James Glass

End-to-end deep learning language or dialect identification systems operate on the spectrogram or other acoustic feature and directly generate identification scores for each class.…

eess.AS2018

Large-scale Speaker Retrieval on Random Speaker Variability Subspace

Suwon Shon, Younggun Lee, Taesu Kim

This paper describes a fast speaker search system to retrieve segments of the same voice identity in the large-scale data. A recent study shows that Locality Sensitive Hashing (LSH…

eess.AS2018

Unsupervised Representation Learning of Speech for Dialect Identification

Suwon Shon, Wei-Ning Hsu, James Glass

In this paper, we explore the use of a factorized hierarchical variational autoencoder (FHVAE) model to learn an unsupervised latent representation for dialect identification (DID)…

eess.AS2018

Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model

Suwon Shon, Hao Tang, James Glass

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently…