5 citations · 7 across the 7 of their papers we have counts for
3 papers · 1 filter
Audio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASR
Pranay Dighe, Prateeth Nayak, Oggi Rudovic +3
Accurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. on the phone) is critical for achieving naturalistic, engaging, and privacy-centric…
Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised Models
Vineet Garg, Ognjen Rudovic, Pranay Dighe +5
We address the problem of detecting speech directed to a device that does not contain a specific wake-word. Specifically, we focus on audio coming from a touch-based invocation. Mi…
CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations
Vin Sachidananda, Shao-Yen Tseng, Erik Marchi +2
Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-La…