7 papers · 1 filter
Stabilizing Bandits using Regularization: Precise Regret and A Quantitative Central Limit Theorem
Budhaditya Halder, Ishan Sengupta, Koustav Chowdhury +2
Statistical inference with bandit data presents fundamental challenges owing to adaptive sampling, which violates the independence assumptions underlying classical asymptotic theor…
Bandit Simulation for Average Reward Inference
Samya Praharaj, Chih-Yu Chang, Koulik Khamaru +1
Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains…
Stable Thompson Sampling: Valid Inference via Variance Inflation
Budhaditya Halder, Shubhayan Pan, Koulik Khamaru
We consider the problem of statistical inference when the data is collected via a Thompson Sampling-type algorithm. While Thompson Sampling (TS) is known to be both asymptotically…
Efficient Inference after Directionally Stable Adaptive Experiments
Zikai Shen, Houssam Zenati, Nathan Kallus +3
We study inference on scalar-valued pathwise differentiable targets after adaptive data collection, such as a bandit algorithm. We introduce a novel target-specific condition, dire…
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
Samya Praharaj, Koulik Khamaru
Statistical inference in contextual bandits is challenging due to the adaptive, non-i.i.d. nature of the data. A growing body of work shows that classical least-squares inference c…
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
Samya Praharaj, Koulik Khamaru
Statistical inference from data generated by multi-armed bandit (MAB) algorithms is challenging due to their adaptive, non-i.i.d. nature. A classical manifestation is that sample a…