activity
20182026
most citedStreaming Language Identification using Combination of Acoustic Representations and ASR Hypotheses

9 citations · 40 across the 31 of their papers we have counts for

collaborators
Showing 2021Show all

12 papers · 1 filter

cs.CL2021

Context-Aware Transformer Transducer for Speech Recognition

Feng-Ju Chang, Jing Liu, Martin Radfar +4

End-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, t…

cs.CL2021

FANS: Fusing ASR and NLU for on-device SLU

Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann +1

Spoken language understanding (SLU) systems translate voice input commands to semantics which are encoded as an intent and pairs of slot tags and values. Most current SLU systems d…

eess.AS2021

Learning a Neural Diff for Speech Models

Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow

As more speech processing applications execute locally on edge devices, a set of resource constraints must be considered. In this work we address one of these constraints, namely o…

eess.AS2021

Bifocal Neural ASR: Exploiting Keyword Spotting for Inference Optimization

Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow

We present Bifocal RNN-T, a new variant of the Recurrent Neural Network Transducer (RNN-T) architecture designed for improved inference time latency on speech recognition tasks. Th…

eess.AS2021

Amortized Neural Networks for Low-Latency Speech Recognition

Jonathan Macoskey, Grant P. Strimel, Jinru Su +1

We introduce Amortized Neural Networks (AmNets), a compute cost- and latency-aware network architecture particularly well-suited for sequence modeling tasks. We apply AmNets to the…

eess.AS2021

Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio

Gokce Keskin, Minhua Wu, Brian King +5

Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design de…