1 citations · 2 across the 6 of their papers we have counts for
6 papers
Ministral 3
Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116
We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…
Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining
Thiziri Nait Saada, Louis Bethune, Michal Klein +3
Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quali…
A simple proof of almost sure convergence for the largest singular value of a product of Gaussian matrices
Thiziri Nait Saada, Alireza Naderi
Let and consider the product of independent matrices , each with i.i.d. normalised $\math…
CARMIL: Context-Aware Regularization on Multiple Instance Learning models for Whole Slide Images
Thiziri Nait Saada, Valentina Di Proietto, Benoit Schmauch +2
Multiple Instance Learning (MIL) models have proven effective for cancer prognosis from Whole Slide Images. However, the original MIL formulation incorrectly assumes the patches of…
Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian Processes
Thiziri Nait-Saada, Alireza Naderi, Jared Tanner
The infinitely wide neural network has been proven a useful and manageable mathematical model that enables the understanding of many phenomena appearing in deep learning. One examp…
On the Initialisation of Wide Low-Rank Feedforward Neural Networks
Thiziri Nait Saada, Jared Tanner
The edge-of-chaos dynamics of wide randomly initialized low-rank feedforward networks are analyzed. Formulae for the optimal weight and bias variances are extended from the full-ra…