Showing eess.ASShow all
3 papers · 1 filter
eess.AS2024
Noise-robust Speech Separation with Fast Generative Correction
Helin Wang, Jesus Villalba, Laureano Moro-Velazquez +3
Speech separation, the task of isolating multiple speech sources from a mixed audio signal, remains challenging in noisy environments. In this paper, we propose a generative correc…
eess.AS2023
Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning
Saurabhchand Bhati, Jesús Villalba, Laureano Moro-Velazquez +2
Visually grounded speech systems learn from paired images and their spoken captions. Recently, there have been attempts to utilize the visually grounded models trained from images…
eess.AS2023
Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognition
Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez +2
Speech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV). We introduce a simple novel technique called Self-FiLM to…