collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

Ine Gevers, Walter Daelemans

Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified…

cs.CL2026

Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness

Victor De Marez, Luna De Bruyne, Walter Daelemans

Factual sycophancy occurs when a language model abandons a correct, verifiable answer under social pressure. Because a flip occurs only when pressure toward a false answer exceeds…

cs.CL2026

Do You Get the Hint? Benchmarking LLMs on the Board Game Concept

Ine Gevers, Walter Daelemans

Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept,…

cs.CL2025

WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization

Ine Gevers, Victor De Marez, Luna De Bruyne +1

In this study, we take a closer look at how Winograd schema challenges can be used to evaluate common sense reasoning in LLMs. Specifically, we evaluate generative models of differ…

cs.CL2024

Bag of Lies: Robustness in Continuous Pre-training BERT

Ine Gevers, Walter Daelemans

This study aims to acquire more insights into the continuous pre-training phase of BERT regarding entity knowledge, using the COVID-19 pandemic as a case study. Since the pandemic…