2 papers
cs.SD2026
SAME: A Semantically-Aligned Music Autoencoder
Julian D. Parker, Zach Evans, CJ Carr +4
Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this wo…
cs.SD2026
Stable Audio 3
Zach Evans, Julian D. Parker, Matthew Rice +4
Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of…