14 citations · 22 across the 3 of their papers we have counts for
3 papers
cs.AI2025★ 14 cited
Measuring AI Ability to Complete Long Software Tasks
Thomas Kwa, Ben West, Joel Becker +23
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…
cs.LG2024★ 2 cited
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Hjalmar Wijk, Tao Lin, Joel Becker +20
Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations fo…
cs.LG2021★ 6 cited
Learning Markov State Abstractions for Deep Reinforcement Learning
Cameron Allen, Neev Parikh, Omer Gottesman +1
A fundamental assumption of reinforcement learning in Markov decision processes (MDPs) is that the relevant decision process is, in fact, Markov. However, when MDPs have rich obser…