activity
20212023
most citedLeveraging Redundancy in Multiple Audio Signals for Far-Field Speech Recognition

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2023

Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition

Saumya Y. Sahai, Jing Liu, Thejaswi Muniyappa +8

We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This archite…

cs.CL2023

Dialog act guided contextual adapter for personalized speech recognition

Feng-Ju Chang, Thejaswi Muniyappa, Kanthashree Mysore Sathyendra +3

Personalization in multi-turn dialogs has been a long standing challenge for end-to-end automatic speech recognition (E2E ASR) models. Recent work on contextual adapters has tackle…

eess.AS20231 cited

Leveraging Redundancy in Multiple Audio Signals for Far-Field Speech Recognition

Feng-Ju Chang, Anastasios Alexandridis, Rupak Vignesh Swaminathan +6

To achieve robust far-field automatic speech recognition (ASR), existing techniques typically employ an acoustic front end (AFE) cascaded with a neural transducer (NT) ASR model. T…

cs.CL2022

Compute Cost Amortized Transformer for Streaming ASR

Yi Xie, Jonathan Macoskey, Martin Radfar +5

We present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Ou…

cs.CL2021

Attentive Contextual Carryover for Multi-Turn End-to-End Spoken Language Understanding

Kai Wei, Thanh Tran, Feng-Ju Chang +8

Recent years have seen significant advances in end-to-end (E2E) spoken language understanding (SLU) systems, which directly predict intents and slots from spoken audio. While dialo…