7 citations · 9 across the 5 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.AI2026
PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks
Claas Beger, Ryan Yi, Melanie Mitchell
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluati…
cs.AI2026
Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks
Adrien Deliège, Claas Beger, Marc Van Droogenbroeck +1
The Abstraction and Reasoning Corpus and related benchmarks evaluate whether AI models can solve novel reasoning tasks, but often leave unclear whether success reflects inference o…