7 citations · 7 across the 2 of their papers we have counts for
4 papers
The best of both worlds: stochastic and adversarial episodic MDPs with unknown transition
Tiancheng Jin, Longbo Huang, Haipeng Luo
We consider the best-of-both-worlds problem for learning an episodic Markov Decision Process through episodes, with the goal of achieving re…
Simultaneously Learning Stochastic and Adversarial Episodic MDPs with Known Transition
Tiancheng Jin, Haipeng Luo
This work studies the problem of learning episodic Markov Decision Processes with known transition and bandit feedback. We develop the first algorithm with a ``best-of-both-worlds'…
Learning Adversarial MDPs with Bandit Feedback and Unknown Transition
Chi Jin, Tiancheng Jin, Haipeng Luo +2
We consider the problem of learning in episodic finite-horizon Markov decision processes with an unknown transition function, bandit feedback, and adversarial losses. We propose an…
Deep Reinforcement Learning for Multi-Driver Vehicle Dispatching and Repositioning Problem
John Holler, Risto Vuorio, Zhiwei Qin +6
Order dispatching and driver repositioning (also known as fleet management) in the face of spatially and temporally varying supply and demand are central to a ride-sharing platform…