4 papers
TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models
Jinho Choo, JunSeung Lee, Jimyeong Kim +3
Large language models (LLMs) demonstrate strong multilingual capabilities, yet often fail to consistently generate responses in the intended language, exhibiting a phenomenon known…
Devstral: Fine-tuning Language Models for Coding Agent Applications
Abhinav Rastogi, Adam Yang, Albert Q. Jiang +100
We introduce Devstral-Small, a lightweight open source model for code agents with the best performance among models below 100B size. In this technical report, we give an overview o…
Voxtral
Alexander H. Liu, Andy Ehrenberg, Andy Lo +103
We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art perfo…
Magistral
Mistral-AI, :, Abhinav Rastogi +98
We introduce Magistral, Mistral's first reasoning model and our own scalable reinforcement learning (RL) pipeline. Instead of relying on existing implementations and RL traces dist…