8 papers · 1 filter
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
Haokun Tian, Stefan Lattner, Charalampos Saitis
Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scal…
Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
Ben Hayes, Charalampos Saitis, György Fazekas
Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We…
Mamba-Diffusion Model with Learnable Wavelet for Controllable Symbolic Music Generation
Jincheng Zhang, György Fazekas, Charalampos Saitis
The recent surge in the popularity of diffusion models for image synthesis has attracted new attention to their potential for generation tasks in other domains. However, their appl…
Designing Neural Synthesizers for Low-Latency Interaction
Franco Caspe, Jordie Shier, Mark Sandler +2
Neural Audio Synthesis (NAS) models offer interactive musical control over high-quality, expressive audio generators. While these models can operate in real-time, they often suffer…
Hybrid Losses for Hierarchical Embedding Learning
Haokun Tian, Stefan Lattner, Brian McFee +1
In traditional supervised learning, the cross-entropy loss treats all incorrect predictions equally, ignoring the relevance or proximity of wrong labels to the correct answer. By l…
Foundation Models for Music: A Survey
Yinghao Ma, Anders Ãland, Anton Ragni +39
In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This com…