1 citations · 3 across the 4 of their papers we have counts for
3 papers · 1 filter
Tight Memory-Regret Lower Bounds for Streaming Bandits
Shaoang Li, Lan Zhang, Junhao Wang +1
In this paper, we investigate the streaming bandits problem, wherein the learner aims to minimize regret by dealing with online arriving arms and sublinear arm memory. We establish…
oIRL: Robust Adversarial Inverse Reinforcement Learning with Temporally Extended Actions
David Venuto, Jhelum Chakravorty, Leonard Boussioux +3
Explicit engineering of reward functions for given environments has been a major hindrance to reinforcement learning methods. While Inverse Reinforcement Learning (IRL) is a soluti…
Avoidance Learning Using Observational Reinforcement Learning
David Venuto, Leonard Boussioux, Junhao Wang +4
Imitation learning seeks to learn an expert policy from sampled demonstrations. However, in the real world, it is often difficult to find a perfect expert and avoiding dangerous be…