2 papers
cs.LG2025
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
Ziyi Chen, Junyi Li, Peiran Yu +1
Reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) are important techniques to align large language models (LLM) with human preference. Howe…
math.OC2024
Provably Faster Algorithms for Bilevel Optimization via Without-Replacement Sampling
Junyi Li, Heng Huang
Bilevel Optimization has experienced significant advancements recently with the introduction of new efficient algorithms. Mirroring the success in single-level optimization, stocha…