3 papers
cs.LG2026
Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining
Thiziri Nait Saada, Louis Bethune, Michal Klein +3
Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quali…
cs.CL2026
Ministral 3
Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116
We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…
math.PR2025
A simple proof of almost sure convergence for the largest singular value of a product of Gaussian matrices
Thiziri Nait Saada, Alireza Naderi
Let and consider the product of independent matrices , each with i.i.d. normalised $\math…