activity
20202022
most citedDynamic curriculum learning via data parameters for noise robust keyword spotting

2 citations · 3 across the 5 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2022

Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised Models

Vineet Garg, Ognjen Rudovic, Pranay Dighe +5

We address the problem of detecting speech directed to a device that does not contain a specific wake-word. Specifically, we focus on audio coming from a touch-based invocation. Mi…

eess.AS2021

Multi-task Learning with Cross Attention for Keyword Spotting

Takuya Higuchi, Anmol Gupta, Chandra Dhir

Keyword spotting (KWS) is an important technique for speech applications, which enables users to activate devices by speaking a keyword phrase. Although a phoneme classifier can be…

eess.AS2021

Streaming Transformer for Hardware Efficient Voice Trigger Detection and False Trigger Mitigation

Vineet Garg, Wonil Chang, Siddharth Sigtia +4

We present a unified and hardware efficient architecture for two stage voice trigger detection (VTD) and false trigger mitigation (FTM) tasks. Two stage VTD systems of voice assist…

eess.AS20212 cited

Dynamic curriculum learning via data parameters for noise robust keyword spotting

Takuya Higuchi, Shreyas Saxena, Mehrez Souden +3

We propose dynamic curriculum learning via data parameters for noise robust keyword spotting. Data parameter learning has recently been introduced for image processing, where weigh…

eess.AS20201 cited

Stacked 1D convolutional networks for end-to-end small footprint voice trigger detection

Takuya Higuchi, Mohammad Ghasemzadeh, Kisun You +1

We propose a stacked 1D convolutional neural network (S1DCNN) for end-to-end small footprint voice trigger detection in a streaming scenario. Voice trigger detection is an importan…

eess.AS2020

Hybrid Transformer/CTC Networks for Hardware Efficient Voice Triggering

Saurabh Adya, Vineet Garg, Siddharth Sigtia +2

We consider the design of two-pass voice trigger detection systems. We focus on the networks in the second pass that are used to re-score candidate segments obtained from the first…