From the 1 of 5 linked papers with an AI index.
5 papers
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Haseeb Shah, Lingwei Zhu, Adam White +1
The paper empirically evaluates how design choices in actor‑critic reinforcement learning algorithms affect performance and robustness on a real‑world water‑treatment control task,…
TIFO: Time-Invariant Frequency Operator for Stationarity-Aware Representation Learning in Time Series
Xihao Piao, Zheng Chen, Lingwei Zhu +3
Nonstationary time series forecasting suffers from the distribution shift issue due to the different distributions that produce the training and test data. Existing methods attempt…
Symmetric Behavior Regularized Policy Optimization
Lingwei Zhu, Haseeb Shah, Zheng Chen +2
Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning. This paper is the f…
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
Lingwei Zhu, Han Wang, Yukie Nagai
Sparse continuous policies are distributions that can choose some actions at random yet keep strictly zero probability for the other actions, which are radically different from the…
q-exponential family for policy optimization
Lingwei Zhu, Haseeb Shah, Han Wang +2
Policy optimization methods benefit from a simple and tractable policy parametrization, usually the Gaussian for continuous action spaces. In this paper, we consider a broader poli…