3 papers
cs.LG2026
How Neural Losses Shape VAE Latents
Giorgio Strano, Luca Cerovaz, Michele Mancusi +2
Modern VAEs are rarely trained with the pointwise likelihood implied by the standard -VAE objective. In practice, pointwise reconstruction is often combined with perceptual and…
cs.SD2026
PHALAR: Phasors for Learned Musical Audio Representations
Davide Marincione, Michele Mancusi, Giorgio Strano +4
Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a…
cs.SD2026
EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
Luca Cerovaz, Michele Mancusi, Emanuele RodolÃ
Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated tow…