Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Mislav Balunović, Jasper Dekoninck, Ivo Petrov +2
The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used…
cs.AI2025
ToolFuzz -- Automated Agent Tool Testing
Ivan Milev, Mislav Balunović, Maximilian Baader +1
Large Language Model (LLM) Agents leverage the advanced reasoning capabilities of LLMs in real-world applications. To interface with an environment, these agents often rely on tool…
cs.AI2025
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
Mislav Balunović, Jasper Dekoninck, Nikola Jovanović +2
While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed…