6 papers
Towards End-to-End Automation of AI Research
Yutaro Yamada, Robert Tjarko Lange, Cong Lu +5
The automation of science is a long-standing ambition in the field of AI. While the community has made significant progress in automating individual components of the scientific pr…
Can Learned Optimization Make Reinforcement Learning Less Difficult?
Alexander David Goldie, Chris Lu, Matthew Thomas Jackson +2
While reinforcement learning (RL) holds great potential for decision making in the real world, it suffers from a number of unique difficulties which often need specific considerati…
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu +5
AI is increasingly playing a pivotal role in transforming how scientific discoveries are made. We introduce The AI Scientist-v2, an end-to-end agentic system capable of producing t…
Discovering Preference Optimization Algorithms with and for Large Language Models
Chris Lu, Samuel Holt, Claudio Fanconi +4
Offline preference optimization is a key method for enhancing and controlling the quality of Large Language Model (LLM) outputs. Typically, preference optimization is approached as…
Artificial Generational Intelligence: Cultural Accumulation in Reinforcement Learning
Jonathan Cook, Chris Lu, Edward Hughes +2
Cultural accumulation drives the open-ended and diverse progress in capabilities spanning human history. It builds an expanding body of knowledge and skills by combining individual…
Recurrent Reinforcement Learning with Memoroids
Steven Morad, Chris Lu, Ryan Kortvelesy +3
Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov sta…