2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2026★ 1 cited
Ministral 3
Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116
We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…
cs.SE2025★ 2 cited
Devstral: Fine-tuning Language Models for Coding Agent Applications
Abhinav Rastogi, Adam Yang, Albert Q. Jiang +100
We introduce Devstral-Small, a lightweight open source model for code agents with the best performance among models below 100B size. In this technical report, we give an overview o…
cs.LG2024
Efficient Sparse Training with Structured Dropout
Andy Lo
Dropout is a common regularisation technique in deep learning that improves generalisation. Even though it introduces sparsity and thus potential for higher throughput, it usually…