12 citations · 12 across the 1 of their papers we have counts for
2 papers
stat.ML2019★ 12 cited
An Optimal Private Stochastic-MAB Algorithm Based on an Optimal Private Stopping Rule
Touqir Sajed, Or Sheffet
We present a provably optimal differentially private algorithm for the stochastic multi-arm bandit problem, as opposed to the private analogue of the UCB-algorithm [Mishra and Thak…
stat.ML2018
High-confidence error estimates for learned value functions
Touqir Sajed, Wesley Chung, Martha White
Estimating the value function for a fixed policy is a fundamental problem in reinforcement learning. Policy evaluation algorithms---to estimate value functions---continue to be dev…