activity
20192022
most citedTie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding

8 citations · 14 across the 13 of their papers we have counts for

collaborators

14 papers

cs.SD2022

Sub-8-bit quantization for on-device speech recognition: a regularization-free approach

Kai Zhen, Martin Radfar, Hieu Duy Nguyen +3

For on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Amon…

cs.SD2022

ConvRNN-T: Convolutional Augmented Recurrent Neural Network Transducers for Streaming Speech Recognition

Martin Radfar, Rohit Barnwal, Rupak Vignesh Swaminathan +4

The recurrent neural network transducer (RNN-T) is a prominent streaming end-to-end (E2E) ASR technology. In RNN-T, the acoustic encoder commonly consists of stacks of LSTMs. Very…

cs.CL2022

A neural prosody encoder for end-ro-end dialogue act classification

Kai Wei, Dillon Knox, Martin Radfar +6

Dialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful…

cs.CL2022

Multi-task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

Xuandi Fu, Feng-Ju Chang, Martin Radfar +4

End-to-end Spoken Language Understanding (E2E SLU) has attracted increasing interest due to its advantages of joint optimization and low latency when compared to traditionally casc…

cs.CL2021

Context-Aware Transformer Transducer for Speech Recognition

Feng-Ju Chang, Jing Liu, Martin Radfar +4

End-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, t…

cs.SD2021

Speech Emotion Recognition Using Quaternion Convolutional Neural Networks

Aneesh Muppidi, Martin Radfar

Although speech recognition has become a widespread technology, inferring emotion from speech signals still remains a challenge. To address this problem, this paper proposes a quat…