activity
20162021
most citedStreaming End-to-End Bilingual ASR Systems with Joint Language Identification

9 citations · 28 across the 7 of their papers we have counts for

collaborators

16 papers

eess.AS2022

VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition

Jinhan Wang, Xiaosu Tong, Jinxi Guo +2

While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed…

eess.AS2021

Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio

Gokce Keskin, Minhua Wu, Brian King +5

Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design de…

cs.LG2021

SynthASR: Unlocking Synthetic Data for Speech Recognition

Amin Fazel, Wei Yang, Yulan Liu +4

End-to-end (E2E) automatic speech recognition (ASR) models have recently demonstrated superior performance over the traditional hybrid ASR models. Training an E2E ASR model require…

eess.AS2021

Attention-based Neural Beamforming Layers for Multi-channel Speech Recognition

Bhargav Pulugundla, Yang Gao, Brian King +5

Attention-based beamformers have recently been shown to be effective for multi-channel speech recognition. However, they are less capable at capturing local information. In this wo…

eess.AS2021

Wav2vec-C: A Self-supervised Model for Speech Representation Learning

Samik Sadhu, Di He, Che-Wei Huang +6

Wav2vec-C introduces a novel representation learning technique combining elements from wav2vec 2.0 and VQ-VAE. Our model learns to reproduce quantized representations from partiall…

eess.AS20202 cited

REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling

Hu Hu, Xuesong Yang, Zeynab Raeesy +6

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…