7 papers · 1 filter
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
Ognjen, Rudovic, Pranay Dighe +7
Follow-up conversations with virtual assistants (VAs) enable a user to seamlessly interact with a VA without the need to repeatedly invoke it using a keyword (after the first query…
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Satyam Kumar, Sai Srujana Buddi, Utkarsh Oggy Sarawgi +6
Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increa…
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
Avamarie Brueggeman, Takuya Higuchi, Masood Delfarah +2
Noise robustness is a key aspect of successful speech applications. Speech enhancement (SE) has been investigated to improve automatic speech recognition accuracy; however, its eff…
Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised Models
Vineet Garg, Ognjen Rudovic, Pranay Dighe +5
We address the problem of detecting speech directed to a device that does not contain a specific wake-word. Specifically, we focus on audio coming from a touch-based invocation. Mi…
Streaming Transformer for Hardware Efficient Voice Trigger Detection and False Trigger Mitigation
Vineet Garg, Wonil Chang, Siddharth Sigtia +4
We present a unified and hardware efficient architecture for two stage voice trigger detection (VTD) and false trigger mitigation (FTM) tasks. Two stage VTD systems of voice assist…
Progressive Voice Trigger Detection: Accuracy vs Latency
Siddharth Sigtia, John Bridle, Hywel Richards +3
We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phr…