2 papers
cs.LG2025
PILAF: Optimal Human Preference Sampling for Reward Modeling
Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng +2
As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerge…
stat.ML2024
Localized exploration in contextual dynamic pricing achieves dimension-free regret
Jinhang Chai, Yaqi Duan, Jianqing Fan +1
We study the problem of contextual dynamic pricing with a linear demand model. We propose a novel localized exploration-then-commit (LetC) algorithm which starts with a pure explor…