1 paper · 1 filter
Melissa Ailem, Katerina Marazopoulou, Charlotte Siska +1
Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model's average performance across the test pr…