3 citations · 3 across the 2 of their papers we have counts for
7 papers
Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio
Gokce Keskin, Minhua Wu, Brian King +5
Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design de…
Attention-based Neural Beamforming Layers for Multi-channel Speech Recognition
Bhargav Pulugundla, Yang Gao, Brian King +5
Attention-based beamformers have recently been shown to be effective for multi-channel speech recognition. However, they are less capable at capturing local information. In this wo…
Wav2vec-C: A Self-supervised Model for Speech Representation Learning
Samik Sadhu, Di He, Che-Wei Huang +6
Wav2vec-C introduces a novel representation learning technique combining elements from wav2vec 2.0 and VQ-VAE. Our model learns to reproduce quantized representations from partiall…
Robust Multi-channel Speech Recognition using Frequency Aligned Network
Taejin Park, Kenichi Kumatani, Minhua Wu +1
Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling…
Fully Learnable Front-End for Multi-Channel Acoustic Modeling using Semi-Supervised Learning
Sanna Wager, Aparna Khare, Minhua Wu +2
In this work, we investigated the teacher-student training paradigm to train a fully learnable multi-channel acoustic model for far-field automatic speech recognition (ASR). Using…
Multi-channel Acoustic Modeling using Mixed Bitrate OPUS Compression
Aparna Khare, Shiva Sundaram, Minhua Wu
Recent literature has shown that a learned front end with multi-channel audio input can outperform traditional beam-forming algorithms for automatic speech recognition (ASR). In th…