activity
20182026
most cited(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement

Yoshiki Masuyama, Kohei Saijo, Francesco Paissan +6

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configur…

cs.SD2024

Speech dereverberation constrained on room impulse response characteristics

Louis Bahrman, Mathieu Fontaine, Jonathan Le Roux +1

Single-channel speech dereverberation aims at extracting a dry speech signal from a recording affected by the acoustic reflections in a room. However, most current deep learning-ba…

cs.SD2024

GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model

Haocheng Liu, Teysir Baoueb, Mathieu Fontaine +2

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model…

cs.SD2024

SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis

Teysir Baoueb, Haocheng Liu, Mathieu Fontaine +2

Generative adversarial network (GAN) models can synthesize highquality audio signals while ensuring fast sample generation. However, they are difficult to train and are prone to se…

cs.SD2018

Phasebook and Friends: Leveraging Discrete Representations for Source Separation

Jonathan Le Roux, Gordon Wichern, Shinji Watanabe +2

Deep learning based speech enhancement and source separation systems have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling.…