activity
20162026
most citedHorde of Bandits using Gaussian Markov Random Fields

5 citations · 10 across the 23 of their papers we have counts for

collaborators
Showing 2026Show all

5 papers · 1 filter

cs.LG2026

CAT-Flow: Curvature-Adaptive sTeps for Flow Matching

Qinchan Li, Pedro Cisneros-Velarde, Keru Fu +3

Flow Matching has emerged as a leading framework for generative modeling, powering state-of-the-art systems such as FLUX and Stable Diffusion 3.5. However, the iterative nature of…

cs.LG2026

Convergence of Steepest Descent and Adam under Non-Uniform Smoothness

Sharan Vaswani, Yifan Sun, Reza Babanezhad

Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generali…

cs.LG2026

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs

Michael Lu, Max Qiushi Lin, Mo Chen +1

We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture po…

cs.LG2026

Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes

Max Qiushi Lin, Reza Asad, Kevin Tan +3

Although actor-critic methods have been successful in practice, their theoretical analyses have several limitations. Specifically, existing theoretical work either sidesteps the ex…

cs.LG2026

Towards Parameter-Free Temporal Difference Learning

Yunxiang Li, Mark Schmidt, Reza Babanezhad +1

Temporal difference (TD) learning is a fundamental algorithm for estimating value functions in reinforcement learning. Recent finite-time analyses of TD with linear function approx…