collaborators

6 papers

cs.CL2026

Simplifying the Modeling of Arbitrary Conditionals in Natural Language

Yinhan Lu, Eric Elmoznino, Léo Gagnon +3

Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right decoding and conditional likelihood com…

cs.LG2026

Beyond Distribution Sharpening: The Importance of Task Rewards

Sarthak Mittal, Leo Gagnon, Guillaume Lajoie

Frontier models have demonstrated exceptional capabilities following the integration of task-reward-based reinforcement learning (RL) into their training pipelines, enabling system…

cs.LG2026

A Compression Perspective on Simplicity Bias

Tom Marty, Eric Elmoznino, Leo Gagnon +5

Deep neural networks exhibit a simplicity bias, a well-documented tendency to favor simple functions over complex ones. In this work, we cast new light on this phenomenon through t…

cs.LG2025

Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective

Leo Gagnon, Eric Elmoznino, Sarthak Mittal +4

The rapid adaptation ability of auto-regressive foundation models is often attributed to the diversity of their pre-training data. This is because, from a Bayesian standpoint, mini…

cs.LG2025

Does learning the right latent variables necessarily improve in-context learning?

Sarthak Mittal, Eric Elmoznino, Leo Gagnon +4

Large autoregressive models like Transformers can solve tasks through in-context learning (ICL) without learning new weights, suggesting avenues for efficiently solving new tasks.…

cs.LG2025

In-context learning and Occam's razor

Eric Elmoznino, Tom Marty, Tejas Kasetty +5

A central goal of machine learning is generalization. While the No Free Lunch Theorem states that we cannot obtain theoretical guarantees for generalization without further assumpt…