Showing stat.MEShow all
2 papers · 1 filter
stat.ME2026
Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons
Jiachun Li, David Simchi-Levi, Will Wei Sun
Pairwise human-preference platforms such as Chatbot Arena have become central to large language model (LLM) evaluation, yet reliable task-specific ranking remains challenging. Glob…
stat.ME2026
LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency
Jiachun Li, David Simchi-Levi, Will Wei Sun
Large language model (LLM) evaluation platforms increasingly rely on pairwise human judgments. These data are noisy, sparse, and non-uniform, yet leaderboards are reported with lim…