27 citations · 151 across the 29 of their papers we have counts for
17 papers · 1 filter
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
Thomas Schmied, Jörg Bornschein, Jordi Grau-Moya +2
The success of Large Language Models (LLMs) has sparked interest in various agentic applications. A key hypothesis is that LLMs, leveraging common sense and Chain-of-Thought (CoT)…
Imitating Language via Scalable Inverse Reinforcement Learning
Markus Wulfmeier, Michael Bloesch, Nino Vieillard +13
The majority of language model training builds on imitation learning. It covers pretraining, supervised fine-tuning, and affects the starting conditions for reinforcement learning…
Growing Q-Networks: Solving Continuous Control Tasks with Adaptive Control Resolution
Tim Seyde, Peter Werner, Wilko Schwarting +2
Recent reinforcement learning approaches have shown surprisingly strong capabilities of bang-bang policies for solving continuous control benchmarks. The underlying coarse action s…
Foundations for Transfer in Reinforcement Learning: A Taxonomy of Knowledge Modalities
Markus Wulfmeier, Arunkumar Byravan, Sarah Bechtle +2
Contemporary artificial intelligence systems exhibit rapidly growing abilities accompanied by the growth of required resources, expansive datasets and corresponding investments int…
Replay across Experiments: A Natural Extension of Off-Policy RL
Dhruva Tirumala, Thomas Lampe, Jose Enrique Chen +9
Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to exte…
Equivariant Data Augmentation for Generalization in Offline Reinforcement Learning
Cristina Pinneri, Sarah Bechtle, Markus Wulfmeier +4
We present a novel approach to address the challenge of generalization in offline reinforcement learning (RL), where the agent learns from a fixed dataset without any additional in…