4 papers
Binary Rewards and Reinforcement Learning: Fundamental Challenges
Marc Dymetman
Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for improving reasoning in language models, yet models trained with RLVR often suffer from dive…
Exponential families from a single KL identity
Marc Dymetman
Exponential families encompass the distributions central to modern machine learning -- softmax, Gaussians, and Boltzmann distributions -- and underlie the theory of variational inf…
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity
Germán Kruszewski, Pierre Erbacher, Jos Rozen +1
Reinforcement Learning (RL) has become the de facto standard for tuning LLMs to solve tasks involving reasoning. However, growing evidence shows that models trained in such way oft…
Guaranteed Generation from Large Language Models
Minbeom Kim, Thibaut Thonet, Jos Rozen +3
As large language models (LLMs) are increasingly used across various applications, there is a growing need to control text generation to satisfy specific constraints or requirement…