5 papers
Predicting fixed-sample test decisions enables anytime-valid inference
Chris Holmes, Stephen Walker
Statistical hypothesis tests typically use prespecified sample sizes, yet data often arrive sequentially. Interim analyses invalidate classical error guarantees, while existing seq…
A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care
Oliver Normand, Esther Borsi, Mitch Fruin +6
Large language models (LLMs) often match or exceed clinician-level performance on medical benchmarks, yet very few are evaluated on real clinical data or examined beyond headline m…
Targeting relative risk heterogeneity with causal forests
Vik Shirvaikar, Andrea Storås, Xi Lin +1
The identification of heterogeneous treatment effects (HTE) across subgroups is of significant interest in clinical trial analysis. Several state-of-the-art HTE estimation methods,…
A general framework for probabilistic model uncertainty
Vik Shirvaikar, Stephen G. Walker, Chris Holmes
Existing approaches to model uncertainty typically either compare models using a quantitative model selection criterion or evaluate posterior model probabilities having set a prior…
Confidence in the Reasoning of Large Language Models
Yudi Pawitan, Chris Holmes
There is a growing literature on reasoning by large language models (LLMs), but the discussion on the uncertainty in their responses is still lacking. Our aim is to assess the exte…