2 citations · 2 across the 1 of their papers we have counts for
7 papers
A low latency ASR-free end to end spoken language understanding system
Mohamed Mhiri, Samuel Myer, Vikrant Singh Tomar
In recent years, developing a speech understanding system that classifies a waveform to structured data, such as intents and slots, without first transcribing the speech to text ha…
Deep Convolutional Neural Network-based Inverse Filtering Approach for Speech De-reverberation
Hanwook Chung, Vikrant Singh Tomar, Benoit Champagne
In this paper, we introduce a spectral-domain inverse filtering approach for single-channel speech de-reverberation using deep convolutional neural network (CNN). The main goal is…
Speech Model Pre-training for End-to-End Spoken Language Understanding
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto +2
Whereas conventional spoken language understanding (SLU) systems map speech to text, and then text to intent, end-to-end SLU systems map speech directly to intent through a single…
DONUT: CTC-based Query-by-Example Keyword Spotting
Loren Lugosch, Samuel Myer, Vikrant Singh Tomar
Keyword spotting--or wakeword detection--is an essential feature for hands-free operation of modern voice-controlled devices. With such devices becoming ubiquitous, users might wan…
Efficient keyword spotting using time delay neural networks
Samuel Myer, Vikrant Singh Tomar
This paper describes a novel method of live keyword spotting using a two-stage time delay neural network. The model is trained using transfer learning: initial training with phone…
Tone Recognition Using Lifters and CTC
Loren Lugosch, Vikrant Singh Tomar
In this paper, we present a new method for recognizing tones in continuous speech for tonal languages. The method works by converting the speech signal to a cepstrogram, extracting…