2 papers
stat.ML2025
A Statistical Framework for Ranking LLM-Based Chatbots
Siavash Ameli, Siyuan Zhuang, Ion Stoica +1
Large language models (LLMs) have transformed natural language processing, with frameworks like Chatbot Arena providing pioneering platforms for evaluating these models. By facilit…
cs.AI2025
JudgeBench: A Benchmark for Evaluating LLM-based Judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery +5
LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based…