3 papers
cs.LG2026
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
Navdeep Kumar, Tehila Dahan, Lior Cohen +4
We establish an optimal sample complexity of for obtaining an -optimal global policy using a single-timescale actor-critic (AC) algorithm in infinite-horizon discoun…
stat.CO2025
Semiparametric Robust Estimation of Population Location
Ananyabrata Barua, Ayanendranath Basu
Real-world measurements often comprise a dominant signal contaminated by a noisy background. Robustly estimating the dominant signal in practice has been a fundamental statistical…
cs.LG2025
Monotone and Conservative Policy Iteration Beyond the Tabular Case
S. R. Eshwar, Gugan Thoppe, Ananyabrata Barua +2
We introduce Reliable Policy Iteration (RPI) and Conservative RPI (CRPI), variants of Policy Iteration (PI) and Conservative PI (CPI), that retain tabular guarantees under function…