activity
20172022
most citedLinks: A High-Dimensional Online Clustering Method

13 citations · 27 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

16 papers · 1 filter

eess.AS2022

Augmenting Transformer-Transducer Based Speaker Change Detection With Token-Level Training Loss

Guanlong Zhao, Quan Wang, Han Lu +2

In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T base…

eess.AS2022

Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting

Beltrán Labrador, Guanlong Zhao, Ignacio López Moreno +3

In this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the ke…

eess.AS2022

A Universally-Deployable ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement, and Voice Separation

Tom O'Malley, Arun Narayanan, Quan Wang

Recent work has shown that it is possible to train a single model to perform joint acoustic echo cancellation (AEC), speech enhancement, and voice separation, thereby serving as a…

eess.AS2022

Closing the Gap between Single-User and Multi-User VoiceFilter-Lite

Rajeev Rikhye, Quan Wang, Qiao Liang +2

VoiceFilter-Lite is a speaker-conditioned voice separation model that plays a crucial role in improving speech recognition and speaker verification by suppressing overlapping speec…

eess.AS2022

Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech

Quan Wang, Yang Yu, Jason Pelecanos +2

In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry informa…

eess.AS20211 cited

Cross-attention conformer for context modeling in speech enhancement for ASR

Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley +2

This work introduces \emph{cross-attention conformer}, an attention-based architecture for context modeling in speech enhancement. Given that the context information can often be s…