10 papers
Learning to Drive in New Cities Without Human Demonstrations
Zilin Wang, Saeed Rahmani, Daphne Cornelisse +4
While autonomous vehicles have achieved reliable performance within specific operating regions, their deployment to new cities remains costly and slow. A key bottleneck is the need…
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Edan Toledo, Karen Hambardzumyan, Martin Josifoski +22
AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus o…
How Should We Meta-Learn Reinforcement Learning Algorithms?
Alexander David Goldie, Zilin Wang, Jaron Cohen +2
The process of meta-learning algorithms from data, instead of relying on manual design, is growing in popularity as a paradigm for improving the performance of machine learning sys…
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Bingchen Zhao, Despoina Magka, Minqi Jiang +20
Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce exist…
Ad-Hoc Human-AI Coordination Challenge
Tin DizdareviÄ, Ravi Hammond, Tobias Gessler +7
Achieving seamless coordination between AI agents and humans is crucial for real-world applications, yet it remains a significant open challenge. Hanabi is a cooperative card game…
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
Andrei Lupu, Timon Willi, Jakob Foerster
As Large Language Models (LLMs) gain agentic abilities, they will have to navigate complex multi-agent scenarios, interacting with human users and other agents in cooperative and c…