activity
20162022
most citedTransformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss

26 citations · 42 across the 5 of their papers we have counts for

collaborators

9 papers

cs.LG2022

Contrastive Siamese Network for Semi-supervised Speech Recognition

Soheil Khorram, Jaeyoung Kim, Anshuman Tripathi +3

This paper introduces contrastive siamese (c-siam) network, an architecture for leveraging unlabeled acoustic data in speech recognition. c-siam is the first network that extracts…

eess.AS2021

Reducing Streaming ASR Model Delay with Self Alignment

Jaeyoung Kim, Han Lu, Anshuman Tripathi +2

Reducing prediction delay for streaming end-to-end ASR models with minimal performance regression is a challenging problem. Constrained alignment is a well-known existing approach…

cs.SD2020

Transformer Transducer: One Model Unifying Streaming and Non-streaming Speech Recognition

Anshuman Tripathi, Jaeyoung Kim, Qian Zhang +2

In this paper we present a Transformer-Transducer model architecture and a training technique to unify streaming and non-streaming speech recognition models into one model. The mod…

eess.AS2020

A Density Ratio Approach to Language Model Fusion in End-To-End Automatic Speech Recognition

Erik McDermott, Hasim Sak, Ehsan Variani

This article describes a density ratio approach to integrating external Language Models (LMs) into end-to-end models for Automatic Speech Recognition (ASR). Applied to a Recurrent…

eess.AS202026 cited

Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss

Qian Zhang, Han Lu, Hasim Sak +4

In this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system. Transformer computation blocks…

cs.CL20196 cited

Adversarial Training for Multilingual Acoustic Modeling

Ke Hu, Hasim Sak, Hank Liao

Multilingual training has been shown to improve acoustic modeling performance by sharing and transferring knowledge in modeling different languages. Knowledge sharing is usually ac…