6 citations · 6 across the 1 of their papers we have counts for
1 paper
Vanessa Kosoy
Most known regret bounds for reinforcement learning are either episodic or assume an environment without traps. We derive a regret bound without making either assumption, by allowi…