18 citations
- Himanshu Gupta3 profiles3 · h 10
- Micah Carroll2 profiles3 · h 15
- Michael Kirchhof2 profiles3 · h 10
- Orion Weller2 profiles3 · h 22
- Zheng-Xin Yong3 profiles3 · h 18
- Abdallah Galal2 profiles2 · h 1
- Alon Amit2 profiles2 · h 8
- Ankit Singh2 profiles2 · h 3
- Antonella Pinto2 profiles2 · h 2
- Anton Peristyy2 profiles2 · h 2
- Archimedes T. Apronti2 profiles2 · h 1
- Arunim Agarwal2 profiles2 · h 2
- Cornell UniversityUS2 papers
- Dartmouth CollegeUS2 papers
- Durham UniversityGB2 papers
- McGill UniversityCA2 papers
- Princeton UniversityUS2 papers
- Queen Mary University of LondonGB2 papers
- Rice UniversityUS2 papers
- Sorbonne UniversitéFR2 papers
- The Alan Turing InstituteGB2 papers
- The University of Texas at ArlingtonUS2 papers
- University College LondonGB2 papers
- University of CambridgeGB2 papers
3 papers
cs.LG2026★ 5 cited
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema
Mahdi Naser Moghadasi, Faezeh Ghaderi
We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was run. The motivation came from…
cs.AI2026★ 3 cited
Computational Hermeneutics: Evaluating generative AI as a cultural technology
Cody Kommers, Ruth Ahnert, Maria Antoniak +35
Generative AI systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be measured rather than fundament…
cs.LG2026★ 18 cited
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…