activity
20222024
most citedNon-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech

19 citations · 23 across the 9 of their papers we have counts for

collaborators

9 papers

eess.AS2024

Noise-robust Speech Separation with Fast Generative Correction

Helin Wang, Jesus Villalba, Laureano Moro-Velazquez +3

Speech separation, the task of isolating multiple speech sources from a mixed audio signal, remains challenging in noisy environments. In this paper, we propose a generative correc…

eess.AS20231 cited

Improving fairness for spoken language understanding in atypical speech with Text-to-Speech

Helin Wang, Venkatesh Ravichandran, Milind Rao +8

Spoken language understanding (SLU) systems often exhibit suboptimal performance in processing atypical speech, typically caused by neurological conditions and motor impairments. R…

eess.AS2023

Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning

Saurabhchand Bhati, Jesús Villalba, Laureano Moro-Velazquez +2

Visually grounded speech systems learn from paired images and their spoken captions. Recently, there have been attempts to utilize the visually grounded models trained from images…

eess.AS20231 cited

Regularizing Contrastive Predictive Coding for Speech Applications

Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko +2

Self-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduc…

cs.LG20231 cited

Stabilized training of joint energy-based models and their practical applications

Martin Sustek, Samik Sadhu, Lukas Burget +4

The recently proposed Joint Energy-based Model (JEM) interprets discriminatively trained classifier as an energy model, which is also trained as a generative model describ…

eess.AS2023

Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognition

Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez +2

Speech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV). We introduce a simple novel technique called Self-FiLM to…