7 papers
Recommendation Quality and the Concentration of Consumption: Experimental Evidence from Netflix
Guy Aridor, Winston Chou, Nathan Kallus +3
We study an experiment with 8.5 million users on Netflix's recommender system to measure how improvements in recommendation technology affect the set of products that get consumed.…
Evaluating for the long term: Learnings from industry
Leif Sigerson, Tom Cunningham, Winston Chou +22
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and shar…
A Human-Augmenting Agentic Workflow for Observational Causal Inference
Winston Chou, Adrien Alexandre, Lars Olds +2
Data analysis agents are becoming increasingly common tools for applied and scientific research. Yet, for highly specialized tasks such as Observational Causal Inference (OCI), hum…
Blending Proxy Metrics with a North Star
Winston Chou
Proxy metrics are widely used to improve the precision and velocity of online experimentation (aka A/B testing). Although proxies are often motivated by long-term outcomes that the…
Estimating Representative Causal Effects with Double Machine Learning
Apoorva Lal, Winston Chou
Double Machine Learning is widely used to estimate treatment effects from non-experimental data. The "residuals-on-residuals" regression (RORR) is especially popular for its simpli…
The Value of Personalized Recommendations: Evidence from Netflix
Kevin Zielnicki, Guy Aridor, Aurélien Bibaut +3
Personalized recommendation systems shape much of user choice online, yet their targeted nature makes separating out the value of recommendation and the underlying goods challengin…