Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Do Large Language Model Benchmarks Test Reliability?
Joshua Vendrow, Edward Vendrow, Sara Beery +1
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' g…
cs.LG2025
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…