Showing stat.APShow all
2 papers · 1 filter
stat.AP2026
Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments
Zehao Xu, Tony Paek, Kevin O'Sullivan +1
Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments. LLM-based lab…
stat.AP2026
Decision Quality Evaluation Framework at Pinterest
Yuqi Tian, Robert Paine, Attila Dobi +3
Online platforms require robust systems to enforce content safety policies at scale. A critical component of these systems is the ability to evaluate the quality of moderation deci…