13 citations · 27 across the 10 of their papers we have counts for
16 papers · 1 filter
Augmenting Transformer-Transducer Based Speaker Change Detection With Token-Level Training Loss
Guanlong Zhao, Quan Wang, Han Lu +2
In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T base…
Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting
Beltrán Labrador, Guanlong Zhao, Ignacio López Moreno +3
In this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the ke…
A Universally-Deployable ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement, and Voice Separation
Tom O'Malley, Arun Narayanan, Quan Wang
Recent work has shown that it is possible to train a single model to perform joint acoustic echo cancellation (AEC), speech enhancement, and voice separation, thereby serving as a…
Closing the Gap between Single-User and Multi-User VoiceFilter-Lite
Rajeev Rikhye, Quan Wang, Qiao Liang +2
VoiceFilter-Lite is a speaker-conditioned voice separation model that plays a crucial role in improving speech recognition and speaker verification by suppressing overlapping speec…
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech
Quan Wang, Yang Yu, Jason Pelecanos +2
In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry informa…
Cross-attention conformer for context modeling in speech enhancement for ASR
Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley +2
This work introduces \emph{cross-attention conformer}, an attention-based architecture for context modeling in speech enhancement. Given that the context information can often be s…