6 papers · 1 filter
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
William Overman, Mohsen Bayati
Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and increasing the…
Can We Validate Counterfactual Estimations in the Presence of General Network Interference?
Sadegh Shirani, Yuwei Luo, William Overman +2
Randomized experiments have become a cornerstone of evidence-based decision-making in contexts ranging from online platforms to public health. However, in experimental settings wit…
On Aligning Prediction Models with Clinical Experiential Learning: A Prostate Cancer Case Study
Jacqueline J. Vallon, William Overman, Wanqiao Xu +11
Over the past decade, the use of machine learning (ML) models in healthcare applications has rapidly increased. Despite high performance, modern ML models do not always capture pat…
Higher-Order Causal Message Passing for Experimentation with Complex Interference
Mohsen Bayati, Yuwei Luo, William Overman +2
Accurate estimation of treatment effects is essential for decision-making across various scientific fields. This task, however, becomes challenging in areas like social sciences an…
Aligning Model Properties via Conformal Risk Control
William Overman, Jacqueline Jil Vallon, Mohsen Bayati
AI model alignment is crucial due to inadvertent biases in training data and the underspecified machine learning pipeline, where models with excellent test metrics may not meet end…
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
Kihyun Yu, Duksang Lee, William Overman +1
This paper studies the safe reinforcement learning problem formulated as an episodic finite-horizon tabular constrained Markov decision process with an unknown transition kernel an…