activity
20182024
most citedHuman-Timescale Adaptation in an Open-Ended Task Space

22 citations · 92 across the 12 of their papers we have counts for

collaborators

17 papers

cs.LG2024★ 7 cited

Training Language Models to Self-Correct via Reinforcement Learning

Aviral Kumar, Vincent Zhuang, Rishabh Agarwal +15

Self-correction is a highly desirable capability of large language models (LLMs), yet it has consistently been found to be largely ineffective in modern LLMs. Current methods for t…

cs.LG2024★ 3 cited

Open-Endedness is Essential for Artificial Superhuman Intelligence

Edward Hughes, Michael Dennis, Jack Parker-Holder +5

In recent years there has been a tremendous surge in the general capabilities of AI systems, mainly fuelled by training foundation models on internetscale data. Nevertheless, the c…

cs.LG2024★ 10 cited

Many-Shot In-Context Learning

Rishabh Agarwal, Avi Singh, Lei M. Zhang +12

Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expande…

cs.LG2024★ 12 cited

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis, Ashley Edwards +22

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless…

cs.LG2023★ 3 cited

Vision-Language Models as a Source of Rewards

Kate Baumli, Satinder Baveja, Feryal Behbahani +24

Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for bui…

cs.LG2023★ 11 cited

Structured State Space Models for In-Context Reinforcement Learning

Chris Lu, Yannick Schroecker, Albert Gu +4

Structured state space sequence (S4) models have recently achieved state-of-the-art performance on long-range sequence modeling tasks. These models also have fast inference speeds…