1 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Peiyu Li, Xiuxiu Tang, Si Chen +4
Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluat…