7 papers
Best-Arm Identification with Noisy Actuation
Merve Karakas, Osama Hanna, Lin F. Yang +1
In this paper, we consider a multi-armed bandit (MAB) instance and study how to identify the best arm when arm commands are conveyed from a central learner to a distributed agent o…
Near-Optimal Sample Complexity for Online Constrained MDPs
Chang Liu, Yunfan Li, Lin F. Yang
Safety is a fundamental challenge in reinforcement learning (RL), particularly in real-world applications such as autonomous driving, robotics, and healthcare. To address this, Con…
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
Chang Liu, Yiran Zhao, Lawrence Liu +3
Reinforcement learning (RL) has enhanced the capabilities of large language models (LLMs) through reward-driven training. Nevertheless, this process can introduce excessively long…
Distributed Multi-Agent Bandits Over ErdÅs-Rényi Random Networks
Jingyuan Liu, Hao Qiu, Lin Yang +1
We study the distributed multi-agent multi-armed bandit problem with heterogeneous rewards over random communication graphs. Uniquely, at each time step agents communicate over…
On the optimal regret of collaborative personalized linear bandits
Bruce Huang, Ruida Zhou, Lin F. Yang +1
Stochastic linear bandits are a fundamental model for sequential decision making, where an agent selects a vector-valued action and receives a noisy reward with expected value give…
Does Feedback Help in Bandits with Arm Erasures?
Merve Karakas, Osama Hanna, Lin F. Yang +1
We study a distributed multi-armed bandit (MAB) problem over arm erasure channels, motivated by the increasing adoption of MAB algorithms over communication-constrained networks. I…