Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Symmetric Behavior Regularized Policy Optimization
Lingwei Zhu, Haseeb Shah, Zheng Chen +1
Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning. This paper is the f…
cs.LG2025
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
Lingwei Zhu, Zheng Chen, Yukie Nagai +1
This paper adds to the growing literature of reinforcement learning (RL) for healthcare by proposing a novel paradigm: augmenting any predictor with Rule-based RL Layer (RRLL) that…
cs.LG2025
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
Lingwei Zhu, Han Wang, Yukie Nagai
Sparse continuous policies are distributions that can choose some actions at random yet keep strictly zero probability for the other actions, which are radically different from the…