4 papers
Instance-Adaptive Online Multicalibration
Zhiming Huang, Jamie Morgenstern, Aaron Roth +1
We study online multicalibration beyond the worst-case. We give a single, efficient algorithm which dynamically interpolates between benign and worst-case sequences by adaptively r…
Worst-Case Regret Bounds for Combinatorial Thompson Sampling in Sleeping Semi-Bandits
Zhiming Huang, Bingshan Hu, Jianping Pan
We revisit combinatorial Thompson sampling (CTS) for semi-bandits with sleeping arms, where arm availability varies over time and actions must satisfy combinatorial constraints, as…
Connecting Thompson Sampling and UCB: Towards More Efficient Trade-offs Between Privacy and Regret
Bingshan Hu, Zhiming Huang, Tianyue H. Zhang +2
We address differentially private stochastic bandit problems from the angles of exploring the deep connections among Thompson Sampling with Gaussian priors, Gaussian mechanisms, an…
Efficient and Adaptive Posterior Sampling Algorithms for Bandits
Bingshan Hu, Zhiming Huang, Tianyue H. Zhang +2
We study Thompson Sampling-based algorithms for stochastic bandits with bounded rewards. As the existing problem-dependent regret bound for Thompson Sampling with Gaussian priors […