5 papers
Recommendation Quality and the Concentration of Consumption: Experimental Evidence from Netflix
Guy Aridor, Winston Chou, Nathan Kallus +3
We study an experiment with 8.5 million users on Netflix's recommender system to measure how improvements in recommendation technology affect the set of products that get consumed.…
Evaluating for the long term: Learnings from industry
Leif Sigerson, Tom Cunningham, Winston Chou +22
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and shar…
A Human-Augmenting Agentic Workflow for Observational Causal Inference
Winston Chou, Adrien Alexandre, Lars Olds +2
Data analysis agents are becoming increasingly common tools for applied and scientific research. Yet, for highly specialized tasks such as Observational Causal Inference (OCI), hum…
Blending Proxy Metrics with a North Star
Winston Chou
Proxy metrics are widely used to improve the precision and velocity of online experimentation (aka A/B testing). Although proxies are often motivated by long-term outcomes that the…
Evaluating Decision Rules Across Many Weak Experiments
Winston Chou, Colin Gray, Nathan Kallus +2
Technology firms conduct randomized controlled experiments ("A/B tests") to learn which actions to take to improve business outcomes. In firms with mature experimentation platforms…