2 papers
math.OC2026
Towards Weaker Variance Assumptions for Stochastic Optimization
Ahmet Alacaoglu, Yura Malitsky, Stephen J. Wright
We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed…
cs.LG2026
Inference Time Policy Optimization for Offline RL with Differentiable World Models
Rohan Deb, Stephen J. Wright, Arindam Banerjee
Offline Reinforcement Learning (RL) learns optimal policies from fixed datasets, training a policy once and deploying it at inference time without further refinement. Inspired by m…