5 papers
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Haoyu Han, Heng Yang
Policy-gradient methods are widely used in reinforcement learning, yet training often becomes unstable or slows down as learning progresses. We study this phenomenon through the no…
Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics
Heng Yang
We propose a sampling-based framework for finite-horizon trajectory and policy optimization under differentiable dynamics by casting controller design as inference. Specifically, w…
Sparse Variable Projection in Robotic Perception: Exploiting Separable Structure for Efficient Nonlinear Optimization
Alan Papalia, Nikolas Sanderson, Haoyu Han +3
Robotic perception often requires solving large nonlinear least-squares (NLS) problems. While sparsity has been well-exploited to scale solvers, a complementary and underexploited…
Building Rome with Convex Optimization
Haoyu Han, Heng Yang
Global bundle adjustment is made easy by depth prediction and convex optimization. We (i) propose a scaled bundle adjustment (SBA) formulation that lifts 2D keypoint measurements t…
On the Surprising Robustness of Sequential Convex Optimization for Contact-Implicit Motion Planning
Yulin Li, Haoyu Han, Shucheng Kang +2
Contact-implicit motion planning-embedding contact sequencing as implicit complementarity constraints-holds the promise of leveraging continuous optimization to discover new contac…