5 papers
Binary Rewards and Reinforcement Learning: Fundamental Challenges
Marc Dymetman
Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for improving reasoning in language models, yet models trained with RLVR often suffer from dive…
Exponential families from a single KL identity
Marc Dymetman
Exponential families encompass the distributions central to modern machine learning -- softmax, Gaussians, and Boltzmann distributions -- and underlie the theory of variational inf…
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity
Germán Kruszewski, Pierre Erbacher, Jos Rozen +1
Reinforcement Learning (RL) has become the de facto standard for tuning LLMs to solve tasks involving reasoning. However, growing evidence shows that models trained in such way oft…
FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data
Thibaut Thonet, Germán Kruszewski, Jos Rozen +2
LLM-powered conversational assistants are often deployed in a one-size-fits-all manner, which fails to accommodate individual user preferences. Recently, LLM personalization -- tai…
Guaranteed Generation from Large Language Models
Minbeom Kim, Thibaut Thonet, Jos Rozen +3
As large language models (LLMs) are increasingly used across various applications, there is a growing need to control text generation to satisfy specific constraints or requirement…