13 citations · 23 across the 3 of their papers we have counts for
10 papers
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech
Quan Wang, Yang Yu, Jason Pelecanos +2
In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry informa…
Noisy student-teacher training for robust keyword spotting
Hyun-Jin Park, Pai Zhu, Ignacio Lopez Moreno +1
We propose self-training with noisy student-teacher approach for streaming keyword spotting, that can utilize large-scale unlabeled data and aggressive data augmentation. The propo…
SpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System
Roza Chojnacka, Jason Pelecanos, Quan Wang +1
In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of…
VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
Quan Wang, Ignacio Lopez Moreno, Mert Saglam +8
We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speec…
Training Keyword Spotting Models on Non-IID Data with Federated Learning
Andrew Hard, Kurt Partridge, Cameron Nguyen +5
We demonstrate that a production-quality keyword-spotting model can be trained on-device using federated learning and achieve comparable false accept and false reject rates to a ce…
Personal VAD: Speaker-Conditioned Voice Activity Detection
Shaojin Ding, Quan Wang, Shuo-yiin Chang +2
In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming o…