4 papers · 1 filter
Contextual Online Uncertainty-Aware Preference Learning for Human Feedback
Nan Lu, Ethan Lee, Ethan X. Fang +1
Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. In this paper, we propose a…
Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback
Pangpang Liu, Junwei Lu, Will Wei Sun
We study estimation and statistical inference for reward models used in aligning large language models (LLMs). A key component of LLM alignment is reinforcement learning from human…
Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
Yichi Zhang, Alexander Belloni, Ethan X. Fang +2
Motivated by the need for rigorous and scalable evaluation of large language models, we study contextual preference inference for pairwise comparison functionals of context-depende…
Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation
Zebin Wang, Yi Han, Ethan X. Fang +2
We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs h…