3 papers
cs.CL2026
Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
Baran Peters, Gabor Hollbeck, Robert Jakob +1
As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diag…
stat.AP2026
Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments
Zehao Xu, Tony Paek, Kevin O'Sullivan +1
Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments. LLM-based lab…
stat.AP2026
Decision Quality Evaluation Framework at Pinterest
Yuqi Tian, Robert Paine, Attila Dobi +3
Online platforms require robust systems to enforce content safety policies at scale. A critical component of these systems is the ability to evaluate the quality of moderation deci…