13 papers · 1 filter
Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations
Young Hyun Cho, Franz Stoll, Will Wei Sun +2
Unexpected shocks recur in global operations, requiring decision rules that adapt as market and operating conditions change. Many operational systems also have hierarchical structu…
When Should an AI Workflow Release? Always-Valid Inference for Black-Box Generate-Verify Systems
Young Hyun Cho, Will Wei Sun
LLM-enabled AI workflows increasingly produce outputs through iterative generate-evaluate-revise loops. Each iteration can improve the candidate, but it also creates a release deci…
Policy-Aware Design of Large-Scale Factorial Experiments
Xin Wen, Xi Chen, Will Wei Sun +1
Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, message…
Reinforcement Learning from Human Feedback: A Statistical Perspective
Pangpang Liu, Chengchun Shi, Will Wei Sun
Reinforcement learning from human feedback (RLHF) has emerged as a central framework for aligning large language models (LLMs) with human preferences. Despite its practical success…
Privacy-Preserving Reinforcement Learning from Human Feedback via Decoupled Reward Modeling
Young Hyun Cho, Will Wei Sun
Preference-based fine-tuning has become an important component in training large language models, and the data used at this stage may contain sensitive user information. A central…
Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback
Seong Jin Lee, Will Wei Sun, Yufeng Liu
Reinforcement learning from human feedback (RLHF) has become a cornerstone for aligning large language models with human preferences. However, the heterogeneity of human feedback,…