most citedBenchmarks Saturate When The Model Gets Smarter Than The Judge

1 citations · 1 across the 8 of their papers we have counts for

collaborators

11 papers

cs.LG2026

Scalable Classification of Course Information Sheets Using Large Language Models: A Reusable Institutional Method for Academic Quality Assurance

Brecht Verbeken, Joke Van den Broeck, Inge De Cleyn +4

Purpose: Higher education institutions face increasing pressure to audit course designs for generative AI (GenAI) integration. This paper presents an end-to-end method for using la…

cs.CY2026

Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments

Arne Vanhoyweghen, Vincent Holst, Melika Mobini +9

Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressin…

cs.AI2026

Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)

Brecht Verbeken, Brando Vagenende, Marie-Anne Guerry +2

Large Language Models (LLMs) are increasingly used as scientific copilots, but evidence on their role in research-level mathematics remains limited, especially for workflows access…

cs.LG2026

Probing the Trajectories of Reasoning Traces in Large Language Models

Marthe Ballon, Brecht Verbeken, Vincent Ginis +1

Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and de…

cs.LG2026

Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs

Melika Mobini, Vincent Holst, Floriano Tori +2

Large language models are increasingly used to curate bibliographies, raising the question: are their reference lists distinguishable from human ones? We build paired citation grap…

cs.AI20261 cited

Benchmarks Saturate When The Model Gets Smarter Than The Judge

Marthe Ballon, Andres Algaba, Brecht Verbeken +1

Benchmarks are important tools to track progress in the development of Large Language Models (LLMs), yet inaccuracies in datasets and evaluation methods consistently undermine thei…