4 papers
Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks
Ine Gevers, Walter Daelemans
Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified…
Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness
Victor De Marez, Luna De Bruyne, Walter Daelemans
Factual sycophancy occurs when a language model abandons a correct, verifiable answer under social pressure. Because a flip occurs only when pressure toward a false answer exceeds…
Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
Ine Gevers, Walter Daelemans
Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept,…
WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization
Ine Gevers, Victor De Marez, Luna De Bruyne +1
In this study, we take a closer look at how Winograd schema challenges can be used to evaluate common sense reasoning in LLMs. Specifically, we evaluate generative models of differ…