activity
20182026
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 158 across the 31 of their papers we have counts for

collaborators

38 papers

cs.CL2026

Shieldstral

Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +273

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…

cs.RO2026

Robostral Navigate

Abdelaziz Bounhar, Abhijeet Somani, Aditi Kabra +273

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems d…

cs.AI2026

Voxtral TTS

Mistral-AI, :, Alexander H. Liu +186

We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid…

cs.SD2025

Voxtral

Alexander H. Liu, Andy Ehrenberg, Andy Lo +103

We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art perfo…

cs.CL2025★ 2 cited

Magistral

Mistral-AI, :, Abhinav Rastogi +98

We introduce Magistral, Mistral's first reasoning model and our own scalable reinforcement learning (RL) pipeline. Instead of relying on existing implementations and RL traces dist…

cs.AI2024★ 2 cited

Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness

Khyathi Raghavi Chandu, Linjie Li, Anas Awadalla +5

The ability to acknowledge the inevitable uncertainty in their knowledge and reasoning is a prerequisite for AI systems to be truly truthful and reliable. In this paper, we present…