10 citations · 11 across the 3 of their papers we have counts for
6 papers
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech
Quan Wang, Yang Yu, Jason Pelecanos +2
In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry informa…
SpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System
Roza Chojnacka, Jason Pelecanos, Quan Wang +1
In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of…
Synth2Aug: Cross-domain speaker recognition with TTS synthesized speech
Yiling Huang, Yutian Chen, Jason Pelecanos +1
In recent years, Text-To-Speech (TTS) has been used as a data augmentation technique for speech recognition to help complement inadequacies in the training data. Correspondingly, w…
VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
Quan Wang, Ignacio Lopez Moreno, Mert Saglam +8
We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speec…
The IBM Speaker Recognition System: Recent Advances and Error Analysis
Seyed Omid Sadjadi, Jason Pelecanos, Sriram Ganapathy
We present the recent advances along with an error analysis of the IBM speaker recognition system for conversational speech. Some of the key advancements that contribute to our sys…
The IBM 2016 Speaker Recognition System
Seyed Omid Sadjadi, Sriram Ganapathy, Jason W. Pelecanos
In this paper we describe the recent advancements made in the IBM i-vector speaker recognition system for conversational speech. In particular, we identify key techniques that cont…