6 papers
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Maximo Rulli, Maximo Eduardo Rulli, Thomas Vaitses Fontanari +11
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly cond…
How Neural Losses Shape VAE Latents
Giorgio Strano, Luca Cerovaz, Michele Mancusi +2
Modern VAEs are rarely trained with the pointwise likelihood implied by the standard -VAE objective. In practice, pointwise reconstruction is often combined with perceptual and…
Membership and Dataset Inference Attacks on Large Audio Generative Models
Jakub Proboszcz, PaweÅ Kochanski, Karol Korszun +5
Generative audio models, based on diffusion and autoregressive architectures, have advanced rapidly in both quality and expressiveness. This progress, however, raises pressing copy…
LoopGen: Training-Free Loopable Music Generation
Davide Marincione, Giorgio Strano, Donato Crisostomi +2
Loops--short audio segments designed for seamless repetition--are central to many music genres, particularly those rooted in dance and electronic styles. However, current generativ…
STAGE: Stemmed Accompaniment Generation through Prefix-Based Conditioning
Giorgio Strano, Chiara Ballanti, Donato Crisostomi +3
Recent advances in generative models have made it possible to create high-quality, coherent music, with some systems delivering production-level output. Yet, most existing models f…
Activation Patching for Interpretable Steering in Music Generation
Simone Facchiano, Giorgio Strano, Donato Crisostomi +4
Understanding how large audio models represent music, and using that understanding to steer generation, is both challenging and underexplored. Inspired by mechanistic interpretabil…