activity
20242026
collaborators

9 papers

cs.AI2026

MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Mislav Balunović, Jasper Dekoninck, Ivo Petrov +2

The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used…

cs.AI2025

MathConstruct: Challenging LLM Reasoning with Constructive Proofs

Mislav Balunović, Jasper Dekoninck, Nikola Jovanović +2

While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed…

cs.CL2025

Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Ivo Petrov, Jasper Dekoninck, Lyuben Baltadzhiev +5

Recent math benchmarks for large language models (LLMs) such as MathArena indicate that state-of-the-art reasoning models achieve impressive performance on mathematical competition…

cs.AI2025

ToolFuzz -- Automated Agent Tool Testing

Ivan Milev, Mislav Balunović, Maximilian Baader +1

Large Language Model (LLM) Agents leverage the advanced reasoning capabilities of LLMs in real-world applications. To interface with an environment, these agents often rely on tool…

cs.CL2025

COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act

Philipp Guldimann, Alexander Spiridonov, Robin Staab +9

The EU's Artificial Intelligence Act (AI Act) is a significant step towards responsible AI development, but lacks clear technical interpretation, making it difficult to assess mode…

cs.AI2025

Large Language Models are Advanced Anonymizers

Robin Staab, Mark Vero, Mislav Balunović +1

Recent privacy research on large language models (LLMs) has shown that they achieve near-human-level performance at inferring personal data from online texts. With ever-increasing…