2 papers
cs.CL2025
Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
Ine Gevers, Walter Daelemans
Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept,…
cs.CL2025
WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization
Ine Gevers, Victor De Marez, Luna De Bruyne +1
In this study, we take a closer look at how Winograd schema challenges can be used to evaluate common sense reasoning in LLMs. Specifically, we evaluate generative models of differ…