13 citations · 33 across the 14 of their papers we have counts for
6 papers · 2 filters
A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation
Tom O'Malley, Arun Narayanan, Quan Wang +3
We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speec…
Cross-attention conformer for context modeling in speech enhancement for ASR
Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley +2
This work introduces \emph{cross-attention conformer}, an attention-based architecture for context modeling in speech enhancement. Given that the context information can often be s…
Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection
Wei Xia, Han Lu, Quan Wang +4
In this paper, we present a novel speaker diarization system for streaming on-device applications. In this system, we use a transformer transducer to detect the speaker turns, repr…
Multi-user VoiceFilter-Lite via Attentive Speaker Embedding
Rajeev Rikhye, Quan Wang, Qiao Liang +2
In this paper, we propose a solution to allow speaker conditioned speech models, such as VoiceFilter-Lite, to support an arbitrary number of enrolled users in a single pass. This i…
Personalized Keyphrase Detection using Speaker and Environment Information
Rajeev Rikhye, Quan Wang, Qiao Liang +6
In this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The syst…
SpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System
Roza Chojnacka, Jason Pelecanos, Quan Wang +1
In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of…