4 citations · 9 across the 3 of their papers we have counts for
8 papers
Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
Duc Le, Mahaveer Jain, Gil Keren +9
How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for sp…
Deep Shallow Fusion for RNN-T Personalization
Duc Le, Gil Keren, Julian Chan +3
End-to-end models in general, and Recurrent Neural Network Transducer (RNN-T) in particular, have gained significant traction in the automatic speech recognition community in the l…
Alignment Restricted Streaming Recurrent Neural Network Transducer
Jay Mahadeokar, Yuan Shangguan, Duc Le +6
There is a growing interest in the speech community in developing Recurrent Neural Network Transducer (RNN-T) models for automatic speech recognition (ASR) applications. RNN-T is t…
Contextual RNN-T For Open Domain ASR
Mahaveer Jain, Gil Keren, Jay Mahadeokar +3
End-to-end (E2E) systems for automatic speech recognition (ASR), such as RNN Transducer (RNN-T) and Listen-Attend-Spell (LAS) blend the individual components of a traditional hybri…
N-HANS: Introducing the Augsburg Neuro-Holistic Audio-eNhancement System
Shuo Liu, Gil Keren, Björn Schuller
N-HANS is a Python toolkit for in-the-wild audio enhancement, including speech, music, and general audio denoising, separation, and selective noise or source suppression. The funct…
Single-Channel Speech Separation with Auxiliary Speaker Embeddings
Shuo Liu, Gil Keren, Björn Schuller
We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural n…