activity
20202022
most citedCALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations

5 citations · 6 across the 4 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS20225 cited

CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations

Vin Sachidananda, Shao-Yen Tseng, Erik Marchi +2

Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-La…

eess.AS2021

Analysis and Tuning of a Voice Assistant System for Dysfluent Speech

Vikramjit Mitra, Zifang Huang, Colin Lea +8

Dysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice op…

eess.AS20211 cited

SEP-28k: A Dataset for Stuttering Event Detection From Podcasts With People Who Stutter

Colin Lea, Vikramjit Mitra, Aparna Joshi +2

The ability to automatically detect stuttering events in speech could help speech pathologists track an individual's fluency over time or help improve speech recognition systems fo…

eess.AS2020

Knowledge Transfer for Efficient On-device False Trigger Mitigation

Pranay Dighe, Erik Marchi, Srikanth Vishnubhotla +2

In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as…

eess.AS2020

Detecting Emotion Primitives from Speech and their use in discerning Categorical Emotions

Vasudha Kowtha, Vikramjit Mitra, Chris Bartels +5

Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely…

eess.AS2020

Multi-task Learning for Speaker Verification and Voice Trigger Detection

Siddharth Sigtia, Erik Marchi, Sachin Kajarekar +2

Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single ne…