3 papers
cs.AI2026
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
Lance Ying, Ryan Truong, Prafull Sharma +9
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technolog…
cs.AI2026
Learning Abstractions for Hierarchical Planning in Program-Synthesis Agents
Zergham Ahmed, Kazuki Irie, Joshua B. Tenenbaum +2
Humans learn abstractions and use them to plan efficiently to quickly generalize across tasks -- an ability that remains challenging for state-of-the-art large language model (LLM)…
cs.LG2019
An Evaluation of the Human-Interpretability of Explanation
Isaac Lage, Emily Chen, Jeffrey He +4
Recent years have seen a boom in interest in machine learning systems that can provide a human-understandable rationale for their predictions or decisions. However, exactly what ki…