activity
20222024
most citedA Simple and Provably Efficient Algorithm for Asynchronous Federated Contextual Linear Bandits

3 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.MA2024

Emergency Localization for Mobile Ground Users: An Adaptive UAV Trajectory Planning Method

Zhihao Zhu, Jiafan He, Luyang Hou +3

In emergency search and rescue scenarios, the quick location of trapped people is essential. However, disasters can render the Global Positioning System (GPS) unusable. Unmanned ae…

cs.LG2023

Horizon-free Reinforcement Learning in Adversarial Linear Mixture MDPs

Kaixuan Ji, Qingyue Zhao, Jiafan He +2

Recent studies have shown that episodic reinforcement learning (RL) is no harder than bandits when the total reward is bounded by , and proved regret bounds that have a polyloga…

cs.LG2023

Uniform-PAC Guarantees for Model-Based RL with Bounded Eluder Dimension

Yue Wu, Jiafan He, Quanquan Gu

Recently, there has been remarkable progress in reinforcement learning (RL) with general function approximation. However, all these works only provide regret or sample complexity g…

cs.LG2023

On the Interplay Between Misspecification and Sub-optimality Gap in Linear Contextual Bandits

Weitong Zhang, Jiafan He, Zhiyuan Fan +1

We study linear contextual bandits in the misspecified setting, where the expected reward function can be approximated by a linear function class up to a bounded misspecification l…

cs.LG20231 cited

Variance-Dependent Regret Bounds for Linear Bandits and Reinforcement Learning: Adaptivity and Computational Efficiency

Heyang Zhao, Jiafan He, Dongruo Zhou +2

Recently, several studies (Zhou et al., 2021a; Zhang et al., 2021b; Kim et al., 2021; Zhou and Gu, 2022) have provided variance-dependent regret bounds for linear contextual bandit…

cs.LG20223 cited

A Simple and Provably Efficient Algorithm for Asynchronous Federated Contextual Linear Bandits

Jiafan He, Tianhao Wang, Yifei Min +1

We study federated contextual linear bandits, where agents cooperate with each other to solve a global contextual linear bandit problem with the help of a central server. We co…