From the 1 of 15 linked papers with an AI index.
3 papers · 1 filter
AIMO Interpretability Challenge
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
Brian Rabern, Philipp Mondorf, Barbara Plank
Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master. To address this, we introduce LogicSkil…
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
Philipp Mondorf, Shijia Zhou, Monica Riedler +1
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language models (LLMs) across…