19 citations · 26 across the 3 of their papers we have counts for
10 papers
Attention-based Neural Beamforming Layers for Multi-channel Speech Recognition
Bhargav Pulugundla, Yang Gao, Brian King +5
Attention-based beamformers have recently been shown to be effective for multi-channel speech recognition. However, they are less capable at capturing local information. In this wo…
Wav2vec-C: A Self-supervised Model for Speech Representation Learning
Samik Sadhu, Di He, Che-Wei Huang +6
Wav2vec-C introduces a novel representation learning technique combining elements from wav2vec 2.0 and VQ-VAE. Our model learns to reproduce quantized representations from partiall…
Streaming ResLSTM with Causal Mean Aggregation for Device-Directed Utterance Detection
Xiaosu Tong, Che-Wei Huang, Sri Harish Mallidi +5
In this paper, we propose a streaming model to distinguish voice queries intended for a smart-home device from background speech. The proposed model consists of multiple CNN layers…
Multi-view Frequency LSTM: An Efficient Frontend for Automatic Speech Recognition
Maarten Van Segbroeck, Harish Mallidih, Brian King +3
Acoustic models in real-time speech recognition systems typically stack multiple unidirectional LSTM layers to process the acoustic frames over time. Performance improvements over…
Multi-Stream End-to-End Speech Recognition
Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi +3
Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The…
Multi-encoder multi-resolution framework for end-to-end speech recognition
Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi +3
Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end Automatic Speech Recognition (ASR). The joint…