58 citations · 70 across the 8 of their papers we have counts for
4 papers · 1 filter
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
Tongzhou Mu, Minghua Liu, Hao Su
The success of many RL techniques heavily relies on human-engineered dense rewards, which typically demand substantial domain expertise and extensive trial and error. In our work,…
Constrained Online Two-stage Stochastic Optimization: Algorithm with (and without) Predictions
Piao Hu, Jiashuo Jiang, Guodong Lyu +1
We consider an online two-stage stochastic optimization with long-term constraints over a finite horizon of periods. At each period, we take the first-stage action, observe a m…
Boosting Reinforcement Learning and Planning with Demonstrations: A Survey
Tongzhou Mu, Hao Su
Although reinforcement learning has seen tremendous success recently, this kind of trial-and-error learning can be impractical or inefficient in complex environments. The use of de…
Improving Policy Optimization with Generalist-Specialist Learning
Zhiwei Jia, Xuanlin Li, Zhan Ling +3
Generalization in deep reinforcement learning over unseen environment variations usually requires policy learning over a large set of diverse training variations. We empirically ob…