4 papers
Stabilizing Bandits using Regularization: Precise Regret and A Quantitative Central Limit Theorem
Budhaditya Halder, Ishan Sengupta, Koustav Chowdhury +2
Statistical inference with bandit data presents fundamental challenges owing to adaptive sampling, which violates the independence assumptions underlying classical asymptotic theor…
Bandit Simulation for Average Reward Inference
Samya Praharaj, Chih-Yu Chang, Koulik Khamaru +1
Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains…
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
Samya Praharaj, Koulik Khamaru
Statistical inference in contextual bandits is challenging due to the adaptive, non-i.i.d. nature of the data. A growing body of work shows that classical least-squares inference c…
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
Samya Praharaj, Koulik Khamaru
Statistical inference from data generated by multi-armed bandit (MAB) algorithms is challenging due to their adaptive, non-i.i.d. nature. A classical manifestation is that sample a…