3 papers
cs.LG2026
Beyond Expectations: Learning with Stochastic Dominance Made Practical
Shicong Cen, Jincheng Mei, Hanjun Dai +3
Stochastic dominance serves as a general framework for modeling a broad spectrum of decision preferences under uncertainty, with risk aversion as one notable example, as it natural…
cs.LG2025
Faster WIND: Accelerating Iterative Best-of- Distillation for LLM Alignment
Tong Yang, Jincheng Mei, Hanjun Dai +5
Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algo…
cs.LG2025
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Shicong Cen, Jincheng Mei, Katayoon Goshvadi +6
Reinforcement learning from human feedback (RLHF) has demonstrated great promise in aligning large language models (LLMs) with human preference. Depending on the availability of pr…