activity
20182022
collaborators

11 papers

cs.SD2022

Sub-8-bit quantization for on-device speech recognition: a regularization-free approach

Kai Zhen, Martin Radfar, Hieu Duy Nguyen +3

For on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Amon…

cs.SD2022

ConvRNN-T: Convolutional Augmented Recurrent Neural Network Transducers for Streaming Speech Recognition

Martin Radfar, Rohit Barnwal, Rupak Vignesh Swaminathan +4

The recurrent neural network transducer (RNN-T) is a prominent streaming end-to-end (E2E) ASR technology. In RNN-T, the acoustic encoder commonly consists of stacks of LSTMs. Very…

cs.CL2022

Contextual Adapters for Personalized Speech Recognition in Neural Transducers

Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang +5

Personal rare word recognition in end-to-end Automatic Speech Recognition (E2E ASR) models is a challenge due to the lack of training data. A standard way to address this issue is…

cs.CL2022

A neural prosody encoder for end-ro-end dialogue act classification

Kai Wei, Dillon Knox, Martin Radfar +6

Dialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful…

cs.CL2022

Multi-task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

Xuandi Fu, Feng-Ju Chang, Martin Radfar +4

End-to-end Spoken Language Understanding (E2E SLU) has attracted increasing interest due to its advantages of joint optimization and low latency when compared to traditionally casc…

eess.AS2021

Learning a Neural Diff for Speech Models

Jonathan Macoskey, Grant P. Strimel, Ariya Rastrow

As more speech processing applications execute locally on edge devices, a set of resource constraints must be considered. In this work we address one of these constraints, namely o…