9 papers
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Mislav BalunoviÄ, Jasper Dekoninck, Ivo Petrov +2
The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used…
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
Mislav BalunoviÄ, Jasper Dekoninck, Nikola JovanoviÄ +2
While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed…
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Ivo Petrov, Jasper Dekoninck, Lyuben Baltadzhiev +5
Recent math benchmarks for large language models (LLMs) such as MathArena indicate that state-of-the-art reasoning models achieve impressive performance on mathematical competition…
ToolFuzz -- Automated Agent Tool Testing
Ivan Milev, Mislav BalunoviÄ, Maximilian Baader +1
Large Language Model (LLM) Agents leverage the advanced reasoning capabilities of LLMs in real-world applications. To interface with an environment, these agents often rely on tool…
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
Philipp Guldimann, Alexander Spiridonov, Robin Staab +9
The EU's Artificial Intelligence Act (AI Act) is a significant step towards responsible AI development, but lacks clear technical interpretation, making it difficult to assess mode…
Large Language Models are Advanced Anonymizers
Robin Staab, Mark Vero, Mislav BalunoviÄ +1
Recent privacy research on large language models (LLMs) has shown that they achieve near-human-level performance at inferring personal data from online texts. With ever-increasing…