most citedTrading off rewards and errors in multi-armed bandits

11 citations · 11 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

Stephane Hatgis-Kessell, Emma Brunskill

We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we replace classical RL algorith…

cs.LG2026

PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data

Aishwarya Mandyam, Jason Meng, Ge Gao +4

Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary dataset…

cs.LG2026

Active Learning for Stochastic Contextual Linear Bandits

Emma Brunskill, Ishani Karmarkar, Zhaoqi Li

A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions…

cs.LG202611 cited

Trading off rewards and errors in multi-armed bandits

Akram Erraqabi, Alessandro Lazaric, Michal Valko +2

In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm…

cs.LG2026

Data-driven Error Estimation: Excess Risk Bounds without Class Complexity as Input

Sanath Kumar Krishnamurthy, Anna Lyubarskaja, Emma Brunskill +1

Constructing confidence intervals that are simultaneously valid across a class of estimates is central to tasks such as multiple mean estimation, generalization guarantees, and ada…