most citedMinistral 3

1 citations · 1 across the 2 of their papers we have counts for

collaborators

6 papers

cs.CL20261 cited

Ministral 3

Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116

We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…

cs.LG2025

Learning to Skip the Middle Layers of Transformers

Tim Lawson, Laurence Aitchison

Conditional computation is a popular strategy to make Transformers more efficient. Existing methods often target individual modules (e.g., mixture-of-experts layers) or skip layers…

stat.ML2025

Massively Parallel Expectation Maximization For Approximate Posteriors

Thomas Heap, Sam Bowyer, Laurence Aitchison

Bayesian inference for hierarchical models can be very challenging. MCMC methods have difficulty scaling to large models with many observations and latent variables. While variatio…

cs.AI2025

Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints

Sam Bowyer, Laurence Aitchison, Desi R. Ivanova

Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessm…

cs.LG2025

Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Lucy Farnik, Tim Lawson, Conor Houghton +1

Sparse autoencoders (SAEs) have been successfully used to discover sparse and human-interpretable representations of the latent activations of LLMs. However, we would ultimately li…

cs.LG2025

Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers

Thomas Heap, Tim Lawson, Lucy Farnik +1

Sparse autoencoders (SAEs) are widely used to extract sparse, interpretable latents from transformer activations. We test whether commonly used SAE quality metrics and automatic ex…