26 citations · 42 across the 5 of their papers we have counts for
9 papers
Contrastive Siamese Network for Semi-supervised Speech Recognition
Soheil Khorram, Jaeyoung Kim, Anshuman Tripathi +3
This paper introduces contrastive siamese (c-siam) network, an architecture for leveraging unlabeled acoustic data in speech recognition. c-siam is the first network that extracts…
Reducing Streaming ASR Model Delay with Self Alignment
Jaeyoung Kim, Han Lu, Anshuman Tripathi +2
Reducing prediction delay for streaming end-to-end ASR models with minimal performance regression is a challenging problem. Constrained alignment is a well-known existing approach…
Transformer Transducer: One Model Unifying Streaming and Non-streaming Speech Recognition
Anshuman Tripathi, Jaeyoung Kim, Qian Zhang +2
In this paper we present a Transformer-Transducer model architecture and a training technique to unify streaming and non-streaming speech recognition models into one model. The mod…
A Density Ratio Approach to Language Model Fusion in End-To-End Automatic Speech Recognition
Erik McDermott, Hasim Sak, Ehsan Variani
This article describes a density ratio approach to integrating external Language Models (LMs) into end-to-end models for Automatic Speech Recognition (ASR). Applied to a Recurrent…
Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
Qian Zhang, Han Lu, Hasim Sak +4
In this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system. Transformer computation blocks…
Adversarial Training for Multilingual Acoustic Modeling
Ke Hu, Hasim Sak, Hank Liao
Multilingual training has been shown to improve acoustic modeling performance by sharing and transferring knowledge in modeling different languages. Knowledge sharing is usually ac…