9 papers
Steering dense music retrieval with open-vocabulary concept discovery
Julien Guinot, Alain Riou, Elio Quinton +1
Controllable music retrieval lets users find music that is, for example, more ambient, less distorted, or without guitar while preserving the other semantic content of an original…
On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs
Hendrik Vincent Koops, Hao Hao Tan, Elio Quinton
Recent audio restoration increasingly relies on large-scale conditional latent generative modeling, including diffusion, Schrodinger Bridges, and Flow Matching variants, to invert…
Single-step Controllable Music Bandwidth Extension With Flow Matching
Carlos Hernandez-Olivan, Hendrik Vincent Koops, Hao Hao Tan +1
Audio restoration consists in inverting degradations of a digital audio signal to recover what would have been the pristine quality signal before the degradation occurred. This is…
Not All Deepfakes Are Created Equal: Triaging Audio Forgeries for Robust Deepfake Singer Identification
Davide Salvi, Hendrik Vincent Koops, Elio Quinton
The proliferation of highly realistic singing voice deepfakes presents a significant challenge to protecting artist likeness and content authenticity. Automatic singer identificati…
GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
Julien Guinot, Elio Quinton, György Fazekas
Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area.…
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
Julien Guinot, Alain Riou, Elio Quinton +1
Joint embedding spaces have significantly advanced music understanding and generation by linking text and audio through multimodal contrastive learning. However, these approaches f…