activity
20242026
collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2026

SAME: A Semantically-Aligned Music Autoencoder

Julian D. Parker, Zach Evans, CJ Carr +4

Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this wo…

cs.SD2026

Stable Audio 3

Zach Evans, Julian D. Parker, Matthew Rice +4

Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of…

cs.SD2026

Low-Resource Guidance for Controllable Latent Audio Diffusion

Zachary Novack, Zack Zukowski, CJ Carr +6

Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guid…

cs.SD2025

Fast Text-to-Audio Generation with Adversarial Post-Training

Zachary Novack, Zach Evans, Zack Zukowski +8

Text-to-audio systems, while increasingly performant, are slow at inference time, thus making their latency unpractical for many creative applications. We present Adversarial Relat…

cs.SD2024

Stable Audio Open

Zach Evans, Julian D. Parker, CJ Carr +3

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio mod…

cs.SD2024

Long-form music generation with latent diffusion

Zach Evans, Julian D. Parker, CJ Carr +3

Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text…