17 citations · 24 across the 6 of their papers we have counts for
8 papers · 1 filter
Entropy-Preserving Reinforcement Learning
Aleksei Petrenko, Ben Lipkin, Kevin Chen +4
Policy gradient algorithms have driven many recent advancements in language model reasoning. An appealing property is their ability to learn from exploration on their own trajector…
Reinforcement Learning for Long-Horizon Interactive LLM Agents
Kevin Chen, Marco Cusumano-Towner, Brody Huval +4
Interactive digital agents (IDAs) leverage APIs of stateful digital environments to perform tasks in response to user requests. While IDAs powered by instruction-tuned large langua…
Robust Autonomy Emerges from Self-Play
Marco Cusumano-Towner, David Hafner, Alex Hertzberg +9
Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust an…
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
Sumeet Batra, Bryon Tjanaka, Matthew C. Fontaine +3
Training generally capable agents that thoroughly explore their environment and learn new and diverse skills is a long-term goal of robot learning. Quality Diversity Reinforcement…
Megaverse: Simulating Embodied Agents at One Million Experiences per Second
Aleksei Petrenko, Erik Wijmans, Brennan Shacklett +1
We present Megaverse, a new 3D simulation platform for reinforcement learning and embodied AI research. The efficient design of our engine enables physics-based simulation with hig…
Agents that Listen: High-Throughput Reinforcement Learning with Multiple Sensory Systems
Shashank Hegde, Anssi Kanervisto, Aleksei Petrenko
Humans and other intelligent animals evolved highly sophisticated perception systems that combine multiple sensory modalities. On the other hand, state-of-the-art artificial agents…