18 citations · 56 across the 47 of their papers we have counts for
10 papers · 1 filter
Focus on the present: a regularization method for the ASR source-target attention layer
Nanxin Chen, Piotr Żelasko, Jesús Villalba +1
This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classificati…
Perceptual Loss based Speech Denoising with an ensemble of Audio Pattern Recognition and Self-Supervised Models
Saurabh Kataria, Jesús Villalba, Najim Dehak
Deep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual En…
Learning Speaker Embedding from Text-to-Speech
Jaejin Cho, Piotr Zelasko, Jesus Villalba +2
Zero-shot multi-speaker Text-to-Speech (TTS) generates target speaker voices given an input text and the corresponding speaker embedding. In this work, we investigate the effective…
CopyPaste: An Augmentation Method for Speech Emotion Recognition
Raghavendra Pappagari, Jesús Villalba, Piotr Żelasko +2
Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recogniti…
Self-Expressing Autoencoders for Unsupervised Spoken Term Discovery
Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko +1
Unsupervised spoken term discovery consists of two tasks: finding the acoustic segment boundaries and labeling acoustically similar segments with the same labels. We perform segmen…
Single Channel Far Field Feature Enhancement For Speaker Verification In The Wild
Phani Sankar Nidadavolu, Saurabh Kataria, Paola García-Perera +2
We investigated an enhancement and a domain adaptation approach to make speaker verification systems robust to perturbations of far-field speech. In the enhancement approach, using…