3 papers
cs.LG2026
Learning to Reason Efficiently with Discounted Reinforcement Learning
Alex Ayoub, Kavosh Asadi, Dale Schuurmans +2
Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to…
cs.CL2026
Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size
Dikshant Kukreja, Kshitij Sah, Gautam Gupta +5
Larger language models become simultaneously better and worse at handling contextual information -- better at ignoring false claims, worse at ignoring irrelevant tokens. We formali…
cs.LG2025
Rectifying Regression in Reinforcement Learning
Alex Ayoub, David Szepesvári, Alireza Bakhtiari +2
This paper investigates the impact of the loss function in value-based methods for reinforcement learning through an analysis of underlying prediction objectives. We theoretically…