papers

Publications (6)

cs.CL2023

Practical Conformer: Optimizing size, speed and flops of Conformer for on-Device and cloud ASR

Rami Botros, Anmol Gulati, Tara N. Sainath +5

Conformer models maintain a large number of internal states, the vast majority of which are associated with self-attention layers. With limited memory bandwidth, reading these from…

eess.AS2024

Anatomy of Industrial Scale Multilingual ASR

Francis McCann Ramirez, Luka Chkhetiani, Andrew Ehrenberg +14

This paper describes AssemblyAI's industrial-scale automatic speech recognition (ASR) system, designed to meet the requirements of large-scale, multilingual ASR serving various app…

cs.CL2021

Tied & Reduced RNN-T Decoder

Rami Botros, Tara N. Sainath, Robert David +3

Previous works on the Recurrent Neural Network-Transducer (RNN-T) models have shown that, under some conditions, it is possible to simplify its prediction network with little or no…

eess.AS2019

Connecting and Comparing Language Model Interpolation Techniques

Ernest Pusateri, Christophe Van Gysel, Rami Botros +4

In this work, we uncover a theoretical connection between two language model interpolation techniques, count merging and Bayesian interpolation. We compare these techniques as well…

cs.CL2023

Lego-Features: Exporting modular encoder features for streaming and deliberation ASR

Rami Botros, Rohit Prabhavalkar, Johan Schalkwyk +3

In end-to-end (E2E) speech recognition models, a representational tight-coupling inevitably emerges between the encoder and the decoder. We build upon recent work that has begun to…

eess.AS2022

A Unified Cascaded Encoder ASR Model for Dynamic Model Sizes

Shaojin Ding, Weiran Wang, Ding Zhao +11

In this paper, we propose a dynamic cascaded encoder Automatic Speech Recognition (ASR) model, which unifies models for different deployment scenarios. Moreover, the model can sign…