activity
20182024
most citedP3O: Policy-on Policy-off Policy Optimization

4 citations · 7 across the 3 of their papers we have counts for

collaborators

8 papers

quant-ph2024

AlphaRouter: Quantum Circuit Routing with Reinforcement Learning and Tree Search

Wei Tang, Yiheng Duan, Yaroslav Kharkov +3

Quantum computers have the potential to outperform classical computers in important tasks such as optimization and number factoring. They are characterized by limited connectivity,…

cs.LG20203 cited

DDPG++: Striving for Simplicity in Continuous-control Off-Policy Reinforcement Learning

Rasool Fakoor, Pratik Chaudhari, Alexander J. Smola

This paper prescribes a suite of techniques for off-policy Reinforcement Learning (RL) that simplify the training process and reduce the sample complexity. First, we show that simp…

cs.LG2020

TraDE: Transformers for Density Estimation

Rasool Fakoor, Pratik Chaudhari, Jonas Mueller +1

We present TraDE, a self-attention-based architecture for auto-regressive density estimation with continuous and discrete valued data. Our model is trained using a penalized maximu…

cs.LG2019

Meta-Q-Learning

Rasool Fakoor, Pratik Chaudhari, Stefano Soatto +1

This paper introduces Meta-Q-Learning (MQL), a new off-policy algorithm for meta-Reinforcement Learning (meta-RL). MQL builds upon three simple ideas. First, we show that Q-learnin…

cs.LG20194 cited

P3O: Policy-on Policy-off Policy Optimization

Rasool Fakoor, Pratik Chaudhari, Alexander J. Smola

On-policy reinforcement learning (RL) algorithms have high sample complexity while off-policy algorithms are difficult to tune. Merging the two holds the promise to develop efficie…

cs.LG2018

Differentiable Greedy Networks

Thomas Powers, Rasool Fakoor, Siamak Shakeri +4

Optimal selection of a subset of items from a given set is a hard problem that requires combinatorial optimization. In this paper, we propose a subset selection algorithm that is t…