Publications (18)
The ICSTM+TUM+UP Approach to the 3rd CHIME Challenge: Single-Channel LSTM Speech Enhancement with Multi-Channel Correlation Shaping Dereverberation and LSTM Language Models
Amr El-Desoky Mousa, Erik Marchi, Björn Schuller
This paper presents our contribution to the 3rd CHiME Speech Separation and Recognition Challenge. Our system uses Bidirectional Long Short-Term Memory (BLSTM) Recurrent Neural Net…
Knowledge Transfer for Efficient On-device False Trigger Mitigation
Pranay Dighe, Erik Marchi, Srikanth Vishnubhotla +2
In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as…
CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations
Vin Sachidananda, Shao-Yen Tseng, Erik Marchi +2
Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-La…
Multi-task Learning for Speaker Verification and Voice Trigger Detection
Siddharth Sigtia, Erik Marchi, Sachin Kajarekar +2
Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single ne…
Progressive Voice Trigger Detection: Accuracy vs Latency
Siddharth Sigtia, John Bridle, Hywel Richards +3
We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phr…
Leveraging Acoustic Cues and Paralinguistic Embeddings to Detect Expression from Voice
Vikramjit Mitra, Sue Booker, Erik Marchi +6
Millions of people reach out to digital assistants such as Siri every day, asking for information, making phone calls, seeking assistance, and much more. The expectation is that su…