Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Benchmark^2: Systematic Evaluation of LLM Benchmarks
Qi Qian, Chengsong Huang, Jingwen Xu +13
The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose B…
cs.CL2024
Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods
Polina Tsvilodub, Hening Wang, Sharon Grosch +1
This paper systematically compares different methods of deriving item-level predictions of language models for multiple-choice tasks. It compares scoring methods for answer options…