5 papers
The Costs of Pretending That There Are Data-Generating Probability Distributions in the Social World
Benedikt Höltgen, Robert C. Williamson
Machine Learning research, including work promoting fair or equitable algorithms, often relies on the concept of a data-generating probability distribution. The standard presumptio…
Limits to Predicting Online Speech Using Large Language Models
Mina Remeli, Moritz Hardt, Robert C. Williamson
Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We d…
Sparse Robust Classification via the Kernel Mean
Brendan van Rooyen, Aditya Krishna Menon, Robert C. Williamson
Many leading classification algorithms output a classifier that is a weighted average of kernel evaluations. Optimizing these weights is a nontrivial problem that still attracts mu…
Geometry and Stability of Supervised Learning Problems
Facundo Mémoli, Brantley Vose, Robert C. Williamson
We introduce a notion of distance between supervised learning problems, which we call the Risk distance. This distance, inspired by optimal transport, facilitates stability results…
Formalising causal inference as prediction on a target population
Benedikt Höltgen, Robert C. Williamson
The standard approach to causal modelling especially in social and health sciences is the potential outcomes framework due to Neyman and Rubin. In this framework, observations are…