10 papers
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
William Overman, Mohsen Bayati
Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and increasing the…
Calibrating Conservatism for Scalable Oversight
William Overman, Mohsen Bayati
Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems…
Causal Effects with Unobserved Unit Types in Interacting Human-AI Systems
William Overman, Sadegh Shirani, Mohsen Bayati
We study experiments on interacting populations of humans and AI agents, where both unit types and the interaction network remain unobserved. Although causal effects propagate thro…
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
William Overman, Mohsen Bayati
As increasingly capable agents are deployed, a central safety challenge is how to retain meaningful human control without modifying the underlying system. We study a minimal contro…
Can We Validate Counterfactual Estimations in the Presence of General Network Interference?
Sadegh Shirani, Yuwei Luo, William Overman +2
Randomized experiments have become a cornerstone of evidence-based decision-making in contexts ranging from online platforms to public health. However, in experimental settings wit…
On Aligning Prediction Models with Clinical Experiential Learning: A Prostate Cancer Case Study
Jacqueline J. Vallon, William Overman, Wanqiao Xu +11
Over the past decade, the use of machine learning (ML) models in healthcare applications has rapidly increased. Despite high performance, modern ML models do not always capture pat…