3 papers
cs.LG2026
You Can Learn Tokenization End-to-End with Reinforcement Learning
Sam Dauncey, Roger Wattenhofer
Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend towards architectures becoming increasi…
cs.LG2026
Subliminal Signals in Preference Labels
Isotta Magistrali, Frédéric Berdoz, Sam Dauncey +1
As AI systems approach superhuman capabilities, scalable oversight increasingly relies on LLM-as-a-judge frameworks where models evaluate and guide each other's training. A core as…
cs.LG2025
Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
Ahmad Fraij, Sam Dauncey
Data scarcity drives the need for more sample-efficient large language models. In this work, we use the double descent phenomenon to holistically compare the sample efficiency of d…