22 citations · 48 across the 8 of their papers we have counts for
6 papers · 1 filter
Cliqueformer: Model-Based Optimization with Structured Transformers
Jakub Grudzien Kuba, Pieter Abbeel, Sergey Levine
Large neural networks excel at prediction tasks, but their application to design problems, such as protein engineering or materials discovery, requires solving offline model-based…
Functional Graphical Models: Structure Enables Offline Data-Driven Optimization
Jakub Grudzien Kuba, Masatoshi Uehara, Pieter Abbeel +1
While machine learning models are typically trained to solve prediction problems, we might often want to use them for optimization problems. For example, given a dataset of protein…
IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Philippe Hansen-Estruch, Ilya Kostrikov, Michael Janner +2
Effective offline RL methods require properly handling out-of-distribution actions. Implicit Q-learning (IQL) addresses this by training a Q-function using only dataset actions thr…
Heterogeneous-Agent Reinforcement Learning
Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng +3
The necessity for cooperation among intelligent machines has popularised cooperative multi-agent reinforcement learning (MARL) in AI research. However, many research endeavours hea…
Discovered Policy Optimisation
Chris Lu, Jakub Grudzien Kuba, Alistair Letcher +3
Tremendous progress has been made in reinforcement learning (RL) over the past decade. Most of these advancements came through the continual development of new algorithms, which we…
Understanding Value Decomposition Algorithms in Deep Cooperative Multi-Agent Reinforcement Learning
Zehao Dou, Jakub Grudzien Kuba, Yaodong Yang
Value function decomposition is becoming a popular rule of thumb for scaling up multi-agent reinforcement learning (MARL) in cooperative games. For such a decomposition rule to hol…