From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Haseeb Shah, Lingwei Zhu, Adam White +1
The paper empirically evaluates how design choices in actor‑critic reinforcement learning algorithms affect performance and robustness on a real‑world water‑treatment control task,…
cs.LG2025
Symmetric Behavior Regularized Policy Optimization
Lingwei Zhu, Haseeb Shah, Zheng Chen +2
Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning. This paper is the f…
cs.LG2025
q-exponential family for policy optimization
Lingwei Zhu, Haseeb Shah, Han Wang +2
Policy optimization methods benefit from a simple and tractable policy parametrization, usually the Gaussian for continuous action spaces. In this paper, we consider a broader poli…