collaborators

8 papers

cs.LG2026

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

Shobhita Sundaram, John Quan, Ariel Kwiatkowski +3

RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigate a fundamental question: Can a pretra…

cs.LG2026

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training

Ismail Labiad, Mathurin Videau, Matthieu Kowalski +4

Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation. However, exposing gradients during training can leak se…

cs.LG2026

Learning through Internalization

Nikolaos Tsilivis, Nirmit Joshi, Marko Medvedev +2

We study internalization processes, by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how they facilitate learning. We in…

cs.LG2026

Efficient RL Training for LLMs with Experience Replay

Charles Arnal, Vivien Cabannes, Taco Cohen +2

While Experience Replay - the practice of storing rollouts and reusing them multiple times during training - is a foundational technique in general RL, it remains largely unexplore…

cs.CL2026

Likelihood-Based Reward Designs for General LLM Reasoning

Ariel Kwiatkowski, Natasha Butt, Ismail Labiad +2

Fine-tuning large language models (LLMs) on reasoning benchmarks via reinforcement learning requires a specific reward function, often binary, for each benchmark. This comes with t…

cs.CL2025

Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation

Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner

To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally ine…