1 citations · 1 across the 4 of their papers we have counts for
4 papers
End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders
Jixuan Wang, Martin Radfar, Kai Wei +1
It is challenging to extract semantic meanings directly from audio signals in spoken language understanding (SLU), due to the lack of textual information. Popular end-to-end (E2E)…
Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers
Grant P. Strimel, Yi Xie, Brian King +3
Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures…
Leveraging Redundancy in Multiple Audio Signals for Far-Field Speech Recognition
Feng-Ju Chang, Anastasios Alexandridis, Rupak Vignesh Swaminathan +6
To achieve robust far-field automatic speech recognition (ASR), existing techniques typically employ an acoustic front end (AFE) cascaded with a neural transducer (NT) ASR model. T…
Compute Cost Amortized Transformer for Streaming ASR
Yi Xie, Jonathan Macoskey, Martin Radfar +5
We present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Ou…