4 papers
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
Yang Liu, Li Wan, Yiteng Huang +5
Smart glasses are increasingly positioned as the next-generation interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in re…
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Oggi Rudovic, Sameer Dharur +7
Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performanc…
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Satyam Kumar, Sai Srujana Buddi, Utkarsh Oggy Sarawgi +6
Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increa…
Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
Utkarsh Oggy Sarawgi, John Berkowitz, Vineet Garg +5
Streaming neural network models for fast frame-wise responses to various speech and sensory signals are widely adopted on resource-constrained platforms. Hence, increasing the lear…