35 citations · 41 across the 7 of their papers we have counts for
6 papers · 1 filter
Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio
Gokce Keskin, Minhua Wu, Brian King +5
Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design de…
Scaling Laws for Acoustic Models
Jasha Droppo, Oguz Elibol
There is a recent trend in machine learning to increase model quality by growing models to sizes previously thought to be unreasonable. Recent work has shown that autoregressive ge…
Attention-based Neural Beamforming Layers for Multi-channel Speech Recognition
Bhargav Pulugundla, Yang Gao, Brian King +5
Attention-based beamformers have recently been shown to be effective for multi-channel speech recognition. However, they are less capable at capturing local information. In this wo…
Wav2vec-C: A Self-supervised Model for Speech Representation Learning
Samik Sadhu, Di He, Che-Wei Huang +6
Wav2vec-C introduces a novel representation learning technique combining elements from wav2vec 2.0 and VQ-VAE. Our model learns to reproduce quantized representations from partiall…
Detection of Lexical Stress Errors in Non-Native (L2) English with Data Augmentation and Attention
Daniel Korzekwa, Roberto Barra-Chicote, Szymon Zaporowski +6
This paper describes two novel complementary techniques that improve the detection of lexical stress errors in non-native (L2) English speech: attention-based feature extraction an…
Efficient minimum word error rate training of RNN-Transducer for end-to-end speech recognition
Jinxi Guo, Gautam Tiwari, Jasha Droppo +4
In this work, we propose a novel and efficient minimum word error rate (MWER) training method for RNN-Transducer (RNN-T). Unlike previous work on this topic, which performs on-the-…