most citedEfficient Multimodal Neural Networks for Trigger-less Voice Assistants

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20241 cited

Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Shruti Palaskar, Oggi Rudovic, Sameer Dharur +7

Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performanc…

eess.AS2024

Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness

Satyam Kumar, Sai Srujana Buddi, Utkarsh Oggy Sarawgi +6

Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increa…

cs.SD2023

Modality Dropout for Multimodal Device Directed Speech Detection using Verbal and Non-Verbal Features

Gautam Krishna, Sameer Dharur, Oggi Rudovic +4

Device-directed speech detection (DDSD) is the binary classification task of distinguishing between queries directed at a voice assistant versus side conversation or background spe…

cs.LG2023

eDKM: An Efficient and Accurate Train-time Weight Clustering for Large Language Models

Minsik Cho, Keivan A. Vahid, Qichen Fu +5

Since Large Language Models or LLMs have demonstrated high-quality performance on many complex language tasks, there is a great interest in bringing these LLMs to mobile devices fo…

cs.LG20232 cited

Efficient Multimodal Neural Networks for Trigger-less Voice Assistants

Sai Srujana Buddi, Utkarsh Oggy Sarawgi, Tashweena Heeramun +4

The adoption of multimodal interactions by Voice Assistants (VAs) is growing rapidly to enhance human-computer interactions. Smartwatches have now incorporated trigger-less methods…