collaborators

11 papers

cs.CL2026

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz +1

Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes.…

cs.LG2026

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

Shobhita Sundaram, John Quan, Ariel Kwiatkowski +3

RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigate a fundamental question: Can a pretra…

cs.CV2026

Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

Elouan Gardès, Seung Eun Yi, Kartik Ahuja +6

We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tuning is often ill-suited to th…

cs.LG2026

Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries

Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi +4

Next-token prediction (NTP) has driven the success of large language models (LLMs), but it struggles with long-horizon reasoning, planning, and creative writing, with these limitat…

cs.LG2026

ReasonCACHE: Teaching LLMs To Reason Without Weight Updates

Sharut Gupta, Phillip Isola, Stefanie Jegelka +4

Can Large language models (LLMs) learn to reason without any weight update and only through in-context learning (ICL)? ICL is strikingly sample-efficient, often learning from only…

cs.LG2025

Operationalizing Quantized Disentanglement

Vitoria Barin-Pacela, Kartik Ahuja, Simon Lacoste-Julien +1

Recent theoretical work established the unsupervised identifiability of quantized factors under any diffeomorphism. The theory assumes that quantization thresholds correspond to ax…