3 papers
cs.SD2025
Live Music Models
Lyria Team, Antoine Caillon, Brian McWilliams +33
We introduce a new class of generative models for music called live music models that produce a continuous stream of music in real-time with synchronized user control. We release M…
cs.SD2025
Streaming Generation for Music Accompaniment
Yusong Wu, Mason Wang, Heidi Lei +5
Music generation models can produce high-fidelity coherent accompaniment given complete audio input, but are limited to editing and loop-based workflows. We study real-time audio-t…
eess.AS2025
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Bowen Zhang, Congchao Guo, Geng Yang +17
We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, w…