1 citations · 2 across the 4 of their papers we have counts for
5 papers · 1 filter
Whispered and Lombard Neural Speech Synthesis
Qiong Hu, Tobias Bleisch, Petko Petkov +3
It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user.…
Knowledge Transfer for Efficient On-device False Trigger Mitigation
Pranay Dighe, Erik Marchi, Srikanth Vishnubhotla +2
In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as…
Progressive Voice Trigger Detection: Accuracy vs Latency
Siddharth Sigtia, John Bridle, Hywel Richards +3
We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phr…
Detecting Emotion Primitives from Speech and their use in discerning Categorical Emotions
Vasudha Kowtha, Vikramjit Mitra, Chris Bartels +5
Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely…
Multi-task Learning for Speaker Verification and Voice Trigger Detection
Siddharth Sigtia, Erik Marchi, Sachin Kajarekar +2
Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single ne…