6 papers
Digital Einstein Experience: Fast Text-to-Speech for Conversational AI
Joanna Rownicka, Kilian Sprenkamp, Antonio Tripiana +2
We describe our approach to create and deliver a custom voice for a conversational AI use-case. More specifically, we provide a voice for a Digital Einstein character, to enable hu…
Comparison of Speech Representations for Automatic Quality Estimation in Multi-Speaker Text-to-Speech Synthesis
Jennifer Williams, Joanna Rownicka, Pilar Oplustil +1
We aim to characterize how different speakers contribute to the perceived output quality of multi-speaker Text-to-Speech (TTS) synthesis. We automatically rate the quality of TTS u…
Multi-scale Octave Convolutions for Robust Speech Recognition
Joanna Rownicka, Peter Bell, Steve Renals
We propose a multi-scale octave convolution layer to learn robust speech representations efficiently. Octave convolutions were introduced by Chen et al [1] in the computer vision f…
Embeddings for DNN speaker adaptive training
Joanna Rownicka, Peter Bell, Steve Renals
In this work, we investigate the use of embeddings for speaker-adaptive training of DNNs (DNN-SAT) focusing on a small amount of adaptation data per speaker. DNN-SAT can be viewed…
Speech Replay Detection with x-Vector Attack Embeddings and Spectral Features
Jennifer Williams, Joanna Rownicka
We present our system submission to the ASVspoof 2019 Challenge Physical Access (PA) task. The objective for this challenge was to develop a countermeasure that identifies speech a…
Analyzing deep CNN-based utterance embeddings for acoustic model adaptation
Joanna Rownicka, Peter Bell, Steve Renals
We explore why deep convolutional neural networks (CNNs) with small two-dimensional kernels, primarily used for modeling spatial relations in images, are also effective in speech r…