activity
20162022
most citedOn the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

21 citations · 61 across the 8 of their papers we have counts for

collaborators

15 papers

cs.SD20212 cited

DarkGAN: Exploiting Knowledge Distillation for Comprehensible Audio Synthesis with GANs

Javier Nistal, Stefan Lattner, Gaël Richard

Generative Adversarial Networks (GANs) have achieved excellent audio synthesis quality in the last years. However, making them operable with semantically meaningful controls remain…

cs.LG2021

Probabilistic semi-nonnegative matrix factorization: a Skellam-based framework

Benoit Fuentes, Gaël Richard

We present a new probabilistic model to address semi-nonnegative matrix factorization (SNMF), called Skellam-SNMF. It is a hierarchical generative model consisting of prior compone…

cs.SD2021

VQCPC-GAN: Variable-Length Adversarial Audio Synthesis Using Vector-Quantized Contrastive Predictive Coding

Javier Nistal, Cyran Aouameur, Stefan Lattner +1

Influenced by the field of Computer Vision, Generative Adversarial Networks (GANs) are often adopted for the audio domain using fixed-size two-dimensional spectrogram representatio…

cs.MM2021

Cross-Modal Music-Video Recommendation: A Study of Design Choices

Laure Pretet, Gael Richard, Geoffroy Peeters

In this work, we study music/video cross-modal recommendation, i.e. recommending a music track for a video or vice versa. We rely on a self-supervised learning paradigm to learn fr…

cs.SD2021

Self-Supervised VQ-VAE for One-Shot Music Style Transfer

Ondřej Cífka, Alexey Ozerov, Umut Şimşekli +1

Neural style transfer, allowing to apply the artistic style of one image to another, has become one of the most widely showcased computer vision applications shortly after its intr…

eess.AS2020

Comparing Representations for Audio Synthesis Using Generative Adversarial Networks

Javier Nistal, Stefan Lattner, Gaël Richard

In this paper, we compare different audio signal representations, including the raw audio waveform and a variety of time-frequency representations, for the task of audio synthesis…