activity
20182026
most citedJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena

492 citations · 945 across the 46 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Yuanhao Ban, Tong Xie, Sohyun An +6

Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness ben…

cs.AI2025

Sleep-time Compute: Beyond Inference Scaling at Test-time

Kevin Lin, Charlie Snell, Yu Wang +4

Scaling test-time compute has emerged as a key ingredient for enabling large language models (LLMs) to solve difficult problems, but comes with high latency and inference cost. We…

cs.AI2024

GameArena: Evaluating LLM Reasoning through Live Computer Games

Lanxiang Hu, Qiyu Li, Anze Xie +4

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination a…

cs.AI2024

TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Xiaoxuan Liu, Jongseok Park, Langxiang Hu +10

Large Language Model (LLM) serving systems batch concurrent user requests to achieve efficient serving. However, in real-world deployments, such inter-request parallelism from batc…

cs.AI2024

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Wei-Lin Chiang, Lianmin Zheng, Ying Sheng +8

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To addres…

cs.AI2024★ 4 cited

Fairness in Serving Large Language Models

Ying Sheng, Shiyi Cao, Dacheng Li +5

High-demand LLM inference services (e.g., ChatGPT and BARD) support a wide range of requests from short chat conversations to long document reading. To ensure that all client reque…