4 papers
Voxtral TTS
Mistral-AI, :, Alexander H. Liu +186
We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid…
Voxtral Realtime
Mistral-AI, :, Alexander H. Liu +166
We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adap…
Reward-Forcing: Autoregressive Video Generation with Reward Feedback
Jingran Zhang, Ning Li, Yuanhao Ban +2
While most prior work in video generation relies on bidirectional architectures, recent efforts have sought to adapt these models into autoregressive variants to support near real-…
LoL: Longer than Longer, Scaling Video Generation to Hour
Justin Cui, Jie Wu, Ming Li +6
Recent research in long-form video generation has shifted from bidirectional to autoregressive models, yet these methods commonly suffer from error accumulation and a loss of long-…