Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks
Claas Beger, Ryan Yi, Melanie Mitchell
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluati…
cs.AI2025
Do AI Models Perform Human-like Abstract Reasoning Across Modalities?
Claas Beger, Ryan Yi, Shuhao Fu +5
OpenAI's o3-preview reasoning model exceeded human accuracy on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models recognize and reason with the abstractions the be…