9 citations · 40 across the 31 of their papers we have counts for
12 papers · 1 filter
Context-Aware Transformer Transducer for Speech Recognition
Feng-Ju Chang, Jing Liu, Martin Radfar +4
End-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, t…
FANS: Fusing ASR and NLU for on-device SLU
Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann +1
Spoken language understanding (SLU) systems translate voice input commands to semantics which are encoded as an intent and pairs of slot tags and values. Most current SLU systems d…
Learning a Neural Diff for Speech Models
Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow
As more speech processing applications execute locally on edge devices, a set of resource constraints must be considered. In this work we address one of these constraints, namely o…
Bifocal Neural ASR: Exploiting Keyword Spotting for Inference Optimization
Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow
We present Bifocal RNN-T, a new variant of the Recurrent Neural Network Transducer (RNN-T) architecture designed for improved inference time latency on speech recognition tasks. Th…
Amortized Neural Networks for Low-Latency Speech Recognition
Jonathan Macoskey, Grant P. Strimel, Jinru Su +1
We introduce Amortized Neural Networks (AmNets), a compute cost- and latency-aware network architecture particularly well-suited for sequence modeling tasks. We apply AmNets to the…
Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio
Gokce Keskin, Minhua Wu, Brian King +5
Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design de…