1 citations · 2 across the 14 of their papers we have counts for
6 papers · 1 filter
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
Biel Tura Vecino, Subhadeep Maji, Aravind Varier +11
The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training…
Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers
Grant P. Strimel, Yi Xie, Brian King +3
Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures…
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
Rahul Pandey, Roger Ren, Qi Luo +7
End-to-End (E2E) automatic speech recognition (ASR) systems used in voice assistants often have difficulties recognizing infrequent words personalized to the user, such as names an…
Learning a Neural Diff for Speech Models
Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow
As more speech processing applications execute locally on edge devices, a set of resource constraints must be considered. In this work we address one of these constraints, namely o…
Bifocal Neural ASR: Exploiting Keyword Spotting for Inference Optimization
Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow
We present Bifocal RNN-T, a new variant of the Recurrent Neural Network Transducer (RNN-T) architecture designed for improved inference time latency on speech recognition tasks. Th…
Amortized Neural Networks for Low-Latency Speech Recognition
Jonathan Macoskey, Grant P. Strimel, Jinru Su +1
We introduce Amortized Neural Networks (AmNets), a compute cost- and latency-aware network architecture particularly well-suited for sequence modeling tasks. We apply AmNets to the…