Showing math.OCShow all
3 papers · 1 filter
math.OC2026
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Haoyu Han, Heng Yang
Policy-gradient methods are widely used in reinforcement learning, yet training often becomes unstable or slows down as learning progresses. We study this phenomenon through the no…
math.OC2025
On the Surprising Robustness of Sequential Convex Optimization for Contact-Implicit Motion Planning
Yulin Li, Haoyu Han, Shucheng Kang +2
Contact-implicit motion planning-embedding contact sequencing as implicit complementarity constraints-holds the promise of leveraging continuous optimization to discover new contac…
math.OC2024
On the Nonsmooth Geometry and Neural Approximation of the Optimal Value Function of Infinite-Horizon Pendulum Swing-up
Haoyu Han, Heng Yang
We revisit the inverted pendulum problem with the goal of understanding and computing the true optimal value function. We start with an observation that the true optimal value func…