2 papers
cs.SD2025
Voxtral
Alexander H. Liu, Andy Ehrenberg, Andy Lo +103
We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art perfo…
cs.CV2024
Pixtral 12B
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +39
We introduce Pixtral-12B, a 12--billion-parameter multimodal language model. Pixtral-12B is trained to understand both natural images and documents, achieving leading performance o…