18 citations · 20 across the 11 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)
Brecht Verbeken, Brando Vagenende, Marie-Anne Guerry +2
Large Language Models (LLMs) are increasingly used as scientific copilots, but evidence on their role in research-level mathematics remains limited, especially for workflows access…
cs.AI2026
Benchmarks Saturate When The Model Gets Smarter Than The Judge
Marthe Ballon, Andres Algaba, Brecht Verbeken +1
Benchmarks are important tools to track progress in the development of Large Language Models (LLMs), yet inaccuracies in datasets and evaluation methods consistently undermine thei…