9 citations · 28 across the 7 of their papers we have counts for
16 papers
VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition
Jinhan Wang, Xiaosu Tong, Jinxi Guo +2
While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed…
Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio
Gokce Keskin, Minhua Wu, Brian King +5
Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design de…
SynthASR: Unlocking Synthetic Data for Speech Recognition
Amin Fazel, Wei Yang, Yulan Liu +4
End-to-end (E2E) automatic speech recognition (ASR) models have recently demonstrated superior performance over the traditional hybrid ASR models. Training an E2E ASR model require…
Attention-based Neural Beamforming Layers for Multi-channel Speech Recognition
Bhargav Pulugundla, Yang Gao, Brian King +5
Attention-based beamformers have recently been shown to be effective for multi-channel speech recognition. However, they are less capable at capturing local information. In this wo…
Wav2vec-C: A Self-supervised Model for Speech Representation Learning
Samik Sadhu, Di He, Che-Wei Huang +6
Wav2vec-C introduces a novel representation learning technique combining elements from wav2vec 2.0 and VQ-VAE. Our model learns to reproduce quantized representations from partiall…
REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling
Hu Hu, Xuesong Yang, Zeynab Raeesy +6
Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…