2 papers
cs.LG2026
Wasserstein Proximal Policy Gradient
Zhaoyu Zhu, Shuhan Zhang, Rui Gao +1
We study policy gradient methods for continuous-action, entropy-regularized reinforcement learning through the lens of Wasserstein geometry. Starting from a Wasserstein proximal up…
cs.LG2026
DeepHalo: A Neural Choice Model with Controllable Context Effects
Shuhan Zhang, Zhi Wang, Rui Gao +1
Modeling human decision-making is central to applications such as recommendation, preference learning, and human-AI alignment. While many classic models assume context-independent…