2 citations · 3 across the 5 of their papers we have counts for
5 papers
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Oggi Rudovic, Sameer Dharur +7
Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performanc…
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Satyam Kumar, Sai Srujana Buddi, Utkarsh Oggy Sarawgi +6
Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increa…
Modality Dropout for Multimodal Device Directed Speech Detection using Verbal and Non-Verbal Features
Gautam Krishna, Sameer Dharur, Oggi Rudovic +4
Device-directed speech detection (DDSD) is the binary classification task of distinguishing between queries directed at a voice assistant versus side conversation or background spe…
eDKM: An Efficient and Accurate Train-time Weight Clustering for Large Language Models
Minsik Cho, Keivan A. Vahid, Qichen Fu +5
Since Large Language Models or LLMs have demonstrated high-quality performance on many complex language tasks, there is a great interest in bringing these LLMs to mobile devices fo…
Efficient Multimodal Neural Networks for Trigger-less Voice Assistants
Sai Srujana Buddi, Utkarsh Oggy Sarawgi, Tashweena Heeramun +4
The adoption of multimodal interactions by Voice Assistants (VAs) is growing rapidly to enhance human-computer interactions. Smartwatches have now incorporated trigger-less methods…