Showing stat.MLShow all
3 papers · 1 filter
stat.ML2025
Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback
Pangpang Liu, Junwei Lu, Will Wei Sun
We study estimation and statistical inference for reward models used in aligning large language models (LLMs). A key component of LLM alignment is reinforcement learning from human…
stat.ML2025
Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
Yichi Zhang, Alexander Belloni, Ethan X. Fang +2
Motivated by the need for rigorous and scalable evaluation of large language models, we study contextual preference inference for pairwise comparison functionals of context-depende…
stat.ML2024
Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation
Zebin Wang, Yi Han, Ethan X. Fang +2
We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs h…