2 citations · 3 across the 2 of their papers we have counts for
4 papers
Ministral 3
Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116
We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…
Learning Task-Agnostic Representations through Multi-Teacher Distillation
Philippe Formont, Maxime Darrin, Banafsheh Karimian +5
Casting complex inputs into tractable representations is a critical step across various fields. Diverse embedding models emerge from differences in architectures, loss functions, i…
Devstral: Fine-tuning Language Models for Coding Agent Applications
Abhinav Rastogi, Adam Yang, Albert Q. Jiang +100
We introduce Devstral-Small, a lightweight open source model for code agents with the best performance among models below 100B size. In this technical report, we give an overview o…
Statistical Deficiency for Task Inclusion Estimation
Loïc Fosse, Frédéric Béchet, Benoît Favre +5
Tasks are central in machine learning, as they are the most natural objects to assess the capabilities of current models. The trend is to build general models able to address any t…