activity
20212023
most citedILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale

7 citations · 10 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL2023

Personalized Predictive ASR for Latency Reduction in Voice Assistants

Andreas Schwarz, Di He, Maarten Van Segbroeck +2

Streaming Automatic Speech Recognition (ASR) in voice assistants can utilize prefetching to partially hide the latency of response generation. Prefetching involves passing a prelim…

cs.LG20231 cited

Accelerator-Aware Training for Transducer-Based Speech Recognition

Suhaila M. Shakiah, Rupak Vignesh Swaminathan, Hieu Duy Nguyen +6

Machine learning model weights and activations are represented in full-precision during training. This leads to performance degradation in runtime when deployed on neural network a…

eess.AS2023

Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers

Grant P. Strimel, Yi Xie, Brian King +3

Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures…

cs.SD2023

Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition

Saumya Y. Sahai, Jing Liu, Thejaswi Muniyappa +8

We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This archite…

cs.CL20227 cited

ILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale

Gopinath Chennupati, Milind Rao, Gurpreet Chadha +11

Incremental learning is one paradigm to enable model building and updating at scale with streaming data. For end-to-end automatic speech recognition (ASR) tasks, the absence of hum…

cs.CL2022

Compute Cost Amortized Transformer for Streaming ASR

Yi Xie, Jonathan Macoskey, Martin Radfar +5

We present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Ou…