1 citations · 2 across the 6 of their papers we have counts for
7 papers · 1 filter
Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes
Yu Chen, Yuhao Liu, Jiatai Huang +2
We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic environments and lack adaptivity…
uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs
Yu Chen, Jiatai Huang, Yan Dai +1
In this paper, we present a novel algorithm, uniINF, for the Heavy-Tailed Multi-Armed Bandits (HTMAB) problem, demonstrating robustness and adaptability in both stochastic and adve…
Banker Online Mirror Descent: A Universal Approach for Delayed Online Bandit Learning
Jiatai Huang, Yan Dai, Longbo Huang
We propose Banker Online Mirror Descent (Banker-OMD), a novel framework generalizing the classical Online Mirror Descent (OMD) technique in the online learning literature. The Bank…
RLx2: Training a Sparse Deep Reinforcement Learning Model from Scratch
Yiqin Tan, Pihe Hu, Ling Pan +2
Training deep reinforcement learning (DRL) models usually requires high computation costs. Therefore, compressing DRL models possesses immense potential for training acceleration a…
Adaptive Best-of-Both-Worlds Algorithm for Heavy-Tailed Multi-Armed Bandits
Jiatai Huang, Yan Dai, Longbo Huang
In this paper, we generalize the concept of heavy-tailed multi-armed bandits to adversarial environments, and develop robust best-of-both-worlds algorithms for heavy-tailed multi-a…
Scale-Free Adversarial Multi-Armed Bandit with Arbitrary Feedback Delays
Jiatai Huang, Yan Dai, Longbo Huang
We consider the Scale-Free Adversarial Multi-Armed Bandit (MAB) problem with unrestricted feedback delays. In contrast to the standard assumption that all losses are -bounde…