2 papers
cs.CL2026
Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks
Ine Gevers, Walter Daelemans
Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified…
cs.CL2026
Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
Ine Gevers, Walter Daelemans
Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept,…