activity
20222026
most citedDeep Iterative Phase Retrieval for Ptychography

7 citations · 7 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

12 papers · 1 filter

eess.AS2026

EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase

Renzheng Shi, Simon Welker, Timo Gerkmann +1

The Griffin-Lim algorithm has been a seminal contribution for phase reconstruction from amplitude spectrograms, however, requiring (infinitely) high algorithmic delay. Its low-dela…

eess.AS2025

Diffusion Buffer for Online Generative Speech Enhancement

Bunlong Lay, Rostislav Makarov, Simon Welker +2

Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called…

eess.AS2025

Real-Time Streaming Mel Vocoding with Generative Flow Matching

Simon Welker, Tal Peer, Timo Gerkmann

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on gen…

eess.AS2024

Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech

Danilo de Oliveira, Julius Richter, Jean-Marie Lemercier +2

Diffusion models have found great success in generating high quality, natural samples of speech, but their potential for density estimation for speech has so far remained largely u…

eess.AS2024

Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models

Jean-Marie Lemercier, Eloi Moliner, Simon Welker +2

This paper presents an unsupervised method for single-channel blind dereverberation and room impulse response (RIR) estimation, called BUDDy. The algorithm is rooted in Bayesian po…

eess.AS2024

EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation

Julius Richter, Yi-Chiao Wu, Steven Krenn +5

We release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totaling in 100 hours of cle…