activity
20202023
most citedCombinatorial Pure Exploration of Dueling Bandit

7 citations · 12 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

12 papers · 1 filter

cs.LG2023

Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback

Yu Chen, Yihan Du, Pihe Hu +3

Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that e…

cs.LG2023★ 2 cited

Provably Safe Reinforcement Learning with Step-wise Violation Constraints

Nuoya Xiong, Yihan Du, Longbo Huang

In this paper, we investigate a novel safe reinforcement learning problem with step-wise violation constraints. Our problem differs from existing works in that we consider stricter…

cs.LG2023

Multi-task Representation Learning for Pure Exploration in Linear Bandits

Yihan Du, Longbo Huang, Wen Sun

Despite the recent success of representation learning in sequential decision making, the study of the pure exploration scenario (i.e., identify the best option and minimize the sam…

cs.LG2022★ 1 cited

Dueling Bandits: From Two-dueling to Multi-dueling

Yihan Du, Siwei Wang, Longbo Huang

We study a general multi-dueling bandit problem, where an agent compares multiple options simultaneously and aims to minimize the regret due to selecting suboptimal arms. This sett…

cs.LG2022★ 2 cited

Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst Path

Yihan Du, Siwei Wang, Longbo Huang

In this paper, we study a novel episodic risk-sensitive Reinforcement Learning (RL) problem, named Iterated CVaR RL, which aims to maximize the tail of the reward-to-go at each ste…

cs.LG2022

Branching Reinforcement Learning

Yihan Du, Wei Chen

In this paper, we propose a novel Branching Reinforcement Learning (Branching RL) model, and investigate both Regret Minimization (RM) and Reward-Free Exploration (RFE) metrics for…