3 papers
math.OC2026
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
Toshinori Kitamura, Arnob Ghosh, Alex Ayoub +2
Projected subgradient descent (PSD) has gained popularity for solving robust Markov decision processes (RMDPs) because it applies to a broader class of uncertainty sets than tradit…
cs.LG2025
Rectifying Regression in Reinforcement Learning
Alex Ayoub, David Szepesvári, Alireza Bakhtiari +2
This paper investigates the impact of the loss function in value-based methods for reinforcement learning through an analysis of underlying prediction objectives. We theoretically…
cs.LG2025
Learning to Reason Efficiently with Discounted Reinforcement Learning
Alex Ayoub, Kavosh Asadi, Dale Schuurmans +2
Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to…