45 citations · 61 across the 6 of their papers we have counts for
7 papers · 1 filter
PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in Control
Ruijie Zheng, Ching-An Cheng, Hal Daumé +2
Temporal action abstractions, along with belief state representations, are a powerful knowledge sharing mechanism for sequential decision making. In this work, we propose a novel v…
Survival Instinct in Offline Reinforcement Learning
Anqi Li, Dipendra Misra, Andrey Kolobov +1
We present a novel observation about the behavior of offline reinforcement learning (RL) algorithms: on many benchmark datasets, offline RL can produce well-performing and safe pol…
Improving Offline RL by Blending Heuristics
Sinong Geng, Aldo Pacchiano, Andrey Kolobov +1
We propose Heuristic Blending (HUBL), a simple performance-improving technique for a broad class of offline RL algorithms based on value bootstrapping. HUBL modifies the Bellman op…
Measuring Sample Efficiency and Generalization in Reinforcement Learning Benchmarks: NeurIPS 2020 Procgen Benchmark
Sharada Mohanty, Jyotish Poonganam, Adrien Gaidon +20
The NeurIPS 2020 Procgen Competition was designed as a centralized benchmark with clearly defined tasks for measuring Sample Efficiency and Generalization in Reinforcement Learning…
Policy Improvement via Imitation of Multiple Oracles
Ching-An Cheng, Andrey Kolobov, Alekh Agarwal
Despite its promise, reinforcement learning's real-world adoption has been hampered by the need for costly exploration to learn a good policy. Imitation learning (IL) mitigates thi…
Safe Reinforcement Learning via Curriculum Induction
Matteo Turchetta, Andrey Kolobov, Shital Shah +2
In safety-critical applications, autonomous agents may need to learn in an environment where mistakes can be very costly. In such settings, the agent needs to behave safely not onl…