1 paper
Jinjie Ni, Fuzhao Xue, Xiang Yue +5
Evaluating large language models (LLMs) is challenging. Traditional ground-truth-based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-…