3 papers
cs.LG2026
Learning to Reason Efficiently with Discounted Reinforcement Learning
Alex Ayoub, Kavosh Asadi, Dale Schuurmans +2
Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to…
math.OC2026
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
Toshinori Kitamura, Arnob Ghosh, Alex Ayoub +2
Projected subgradient descent (PSD) has gained popularity for solving robust Markov decision processes (RMDPs) because it applies to a broader class of uncertainty sets than tradit…
cs.LG2025
Rectifying Regression in Reinforcement Learning
Alex Ayoub, David Szepesvári, Alireza Bakhtiari +2
This paper investigates the impact of the loss function in value-based methods for reinforcement learning through an analysis of underlying prediction objectives. We theoretically…