Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Voxtral TTS
Mistral-AI, :, Alexander H. Liu +186
We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid…
cs.AI2026
Voxtral Realtime
Mistral-AI, :, Alexander H. Liu +166
We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adap…
cs.AI2024
From Goal-Conditioned to Language-Conditioned Agents via Vision-Language Models
Theo Cachet, Christopher R. Dance, Olivier Sigaud
Vision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. T…