2 papers
cs.AI2025
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
Zihui Zhao, Zechang Li
Direct Preference Optimization (DPO) has emerged as a lightweight and effective alternative to Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning with AI…
math.AP2025
On the asymptotic properties of solutions to one-phase free boundary problems
Max Engelstein, Daniel Restrepo, Zihui Zhao
In this article we study the structure of solutions to the one-phase Bernoulli problem that are modeled either infinitesimally or at infinity by one-homogeneous solutions with an i…