Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
Allen Nie, Anirudhan Badrinath, Nicholas Tomlin +5
Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain exp…
cs.AI2025
Measuring General Intelligence with Generated Games
Vivek Verma, David Huang, William Chen +2
We present gg-bench, a collection of game environments designed to evaluate general reasoning capabilities in language models. Unlike most static benchmarks, gg-bench is a data gen…
cs.AI2024
Autonomous Evaluation and Refinement of Digital Agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin +3
We show that domain-general automatic evaluators can significantly improve the performance of agents for web navigation and device control. We experiment with multiple evaluation m…