activity
20162020
most citedA low latency ASR-free end to end spoken language understanding system

2 citations · 2 across the 1 of their papers we have counts for

collaborators

7 papers

cs.CV20202 cited

A low latency ASR-free end to end spoken language understanding system

Mohamed Mhiri, Samuel Myer, Vikrant Singh Tomar

In recent years, developing a speech understanding system that classifies a waveform to structured data, such as intents and slots, without first transcribing the speech to text ha…

cs.SD2020

Deep Convolutional Neural Network-based Inverse Filtering Approach for Speech De-reverberation

Hanwook Chung, Vikrant Singh Tomar, Benoit Champagne

In this paper, we introduce a spectral-domain inverse filtering approach for single-channel speech de-reverberation using deep convolutional neural network (CNN). The main goal is…

eess.AS2019

Speech Model Pre-training for End-to-End Spoken Language Understanding

Loren Lugosch, Mirco Ravanelli, Patrick Ignoto +2

Whereas conventional spoken language understanding (SLU) systems map speech to text, and then text to intent, end-to-end SLU systems map speech directly to intent through a single…

cs.LG2018

DONUT: CTC-based Query-by-Example Keyword Spotting

Loren Lugosch, Samuel Myer, Vikrant Singh Tomar

Keyword spotting--or wakeword detection--is an essential feature for hands-free operation of modern voice-controlled devices. With such devices becoming ubiquitous, users might wan…

eess.AS2018

Efficient keyword spotting using time delay neural networks

Samuel Myer, Vikrant Singh Tomar

This paper describes a novel method of live keyword spotting using a two-stage time delay neural network. The model is trained using transfer learning: initial training with phone…

eess.AS2018

Tone Recognition Using Lifters and CTC

Loren Lugosch, Vikrant Singh Tomar

In this paper, we present a new method for recognizing tones in continuous speech for tonal languages. The method works by converting the speech signal to a cepstrogram, extracting…