5 citations · 7 across the 7 of their papers we have counts for
6 papers · 1 filter
Knowledge Transfer for Efficient On-device False Trigger Mitigation
Pranay Dighe, Erik Marchi, Srikanth Vishnubhotla +2
In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as…
Progressive Voice Trigger Detection: Accuracy vs Latency
Siddharth Sigtia, John Bridle, Hywel Richards +3
We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phr…
Generating Multilingual Voices Using Speaker Space Translation Based on Bilingual Speaker Data
Soumi Maiti, Erik Marchi, Alistair Conkie
We present progress towards bilingual Text-to-Speech which is able to transform a monolingual voice to speak a second language while preserving speaker voice quality. We demonstrat…
On the Role of Visual Cues in Audiovisual Speech Enhancement
Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald +4
We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues t…
Detecting Emotion Primitives from Speech and their use in discerning Categorical Emotions
Vasudha Kowtha, Vikramjit Mitra, Chris Bartels +5
Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely…
Multi-task Learning for Speaker Verification and Voice Trigger Detection
Siddharth Sigtia, Erik Marchi, Sachin Kajarekar +2
Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single ne…