collaborators

9 papers

cs.CL2026

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz +1

Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes.…

cs.CL2026

Separating Representation from Reconstruction Enables Scalable Text Encoders

Megi Dervishi, Mathurin Videau, Yann LeCun

While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the r…

cs.LG2026

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training

Ismail Labiad, Mathurin Videau, Matthieu Kowalski +4

Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation. However, exposing gradients during training can leak se…

cs.CL2026

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

João Maria Janeiro, Mathurin Videau, Andrea Caciolai +3

Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes them unreliable. Specifically…

cs.CL2026

Evolutionary Pre-Prompt Optimization for Mathematical Reasoning

Mathurin Videau, Alessandro Leite, Marc Schoenauer +1

Recent advancements have highlighted that large language models (LLMs), when given a small set of task-specific examples, demonstrate remarkable proficiency, a capability that exte…

cs.LG2025

Evolutionary Retrofitting

Mathurin Videau, Mariia Zameshina, Alessandro Leite +3

AfterLearnER (After Learning Evolutionary Retrofitting) consists in applying evolutionary optimization to refine fully trained machine learning models by optimizing a set of carefu…