7 citations · 10 across the 8 of their papers we have counts for
8 papers
Personalized Predictive ASR for Latency Reduction in Voice Assistants
Andreas Schwarz, Di He, Maarten Van Segbroeck +2
Streaming Automatic Speech Recognition (ASR) in voice assistants can utilize prefetching to partially hide the latency of response generation. Prefetching involves passing a prelim…
Accelerator-Aware Training for Transducer-Based Speech Recognition
Suhaila M. Shakiah, Rupak Vignesh Swaminathan, Hieu Duy Nguyen +6
Machine learning model weights and activations are represented in full-precision during training. This leads to performance degradation in runtime when deployed on neural network a…
Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers
Grant P. Strimel, Yi Xie, Brian King +3
Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures…
Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition
Saumya Y. Sahai, Jing Liu, Thejaswi Muniyappa +8
We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This archite…
ILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale
Gopinath Chennupati, Milind Rao, Gurpreet Chadha +11
Incremental learning is one paradigm to enable model building and updating at scale with streaming data. For end-to-end automatic speech recognition (ASR) tasks, the absence of hum…
Compute Cost Amortized Transformer for Streaming ASR
Yi Xie, Jonathan Macoskey, Martin Radfar +5
We present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Ou…