5 citations · 6 across the 4 of their papers we have counts for
8 papers
CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations
Vin Sachidananda, Shao-Yen Tseng, Erik Marchi +2
Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-La…
Streaming on-device detection of device directed speech from voice and touch-based invocation
Ognjen Rudovic, Akanksha Bindal, Vineet Garg +3
When interacting with smart devices such as mobile phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the devic…
Analysis and Tuning of a Voice Assistant System for Dysfluent Speech
Vikramjit Mitra, Zifang Huang, Colin Lea +8
Dysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice op…
SEP-28k: A Dataset for Stuttering Event Detection From Podcasts With People Who Stutter
Colin Lea, Vikramjit Mitra, Aparna Joshi +2
The ability to automatically detect stuttering events in speech could help speech pathologists track an individual's fluency over time or help improve speech recognition systems fo…
Knowledge Transfer for Efficient On-device False Trigger Mitigation
Pranay Dighe, Erik Marchi, Srikanth Vishnubhotla +2
In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as…
On the Role of Visual Cues in Audiovisual Speech Enhancement
Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald +4
We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues t…