7 papers · 1 filter
SAME: A Semantically-Aligned Music Autoencoder
Julian D. Parker, Zach Evans, CJ Carr +4
Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this wo…
Stable Audio 3
Zach Evans, Julian D. Parker, Matthew Rice +4
Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of…
Low-Resource Guidance for Controllable Latent Audio Diffusion
Zachary Novack, Zack Zukowski, CJ Carr +6
Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guid…
Fast Text-to-Audio Generation with Adversarial Post-Training
Zachary Novack, Zach Evans, Zack Zukowski +8
Text-to-audio systems, while increasingly performant, are slow at inference time, thus making their latency unpractical for many creative applications. We present Adversarial Relat…
Stable Audio Open
Zach Evans, Julian D. Parker, CJ Carr +3
Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio mod…
Long-form music generation with latent diffusion
Zach Evans, Julian D. Parker, CJ Carr +3
Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text…