collaborators

5 papers

math.OC2026

Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control

Zhaoyu Zhu, Rui Gao, Shuang Li

Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A B…

cs.LG2026

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

Zhaoyu Zhu, Rui Gao, Shuang Li

Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions. For the entr…

cs.LG2026

C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs

Rui Gao, Youngseung Jeon, Swastik Roy +2

Large language models (LLMs) show promise for molecular optimization, but aligning them with selective and competing drug-design constraints remains challenging. We propose C-Moral…

cs.LG2026

Wasserstein Proximal Policy Gradient

Zhaoyu Zhu, Shuhan Zhang, Rui Gao +1

We study policy gradient methods for continuous-action, entropy-regularized reinforcement learning through the lens of Wasserstein geometry. Starting from a Wasserstein proximal up…

cs.LG2026

DeepHalo: A Neural Choice Model with Controllable Context Effects

Shuhan Zhang, Zhi Wang, Rui Gao +1

Modeling human decision-making is central to applications such as recommendation, preference learning, and human-AI alignment. While many classic models assume context-independent…