2 papers
stat.ML2026
PFN-TS: Thompson Sampling for Contextual Bandits via Prior-Data Fitted Networks
Yan Shuo Tan, Kenyon Ng, Ruizhe Deng +3
Thompson sampling is a widely used strategy for contextual bandits: at each round, it samples a reward function from a Bayesian posterior and acts greedily under that sample. Prior…
stat.ME2025
Bayesian Machine Learning for Estimating Optimal Dynamic Treatment Regimes with Ordinal Outcomes
Xinru Wang, Tanujit Chakraborty, Bibhas Chakraborty
Dynamic treatment regimes (DTRs) are sequences of decision rules designed to tailor treatment based on patients' treatment history and evolving disease status. Ordinal outcomes fre…