4 papers
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Ro Encarnación, Tina Behzad, Emma Lurie +1
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a sing…
The Beginning of ChatGPT Ads
Emma Lurie, Ro Encarnación, Sorelle A. Friedler +1
This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examin…
Triangulating Across U.S. Federal AI Transparency Regimes
Emma Lurie, Emma Fauser, Qing He +2
Federal AI systems can deny benefits or flag individuals for deportation, but the public disclosures meant to make those systems visible are fragmented and unevenly detailed. This…
Longitudinal Monitoring of LLM Content Moderation of Social Issues
Yunlang Dai, Emma Lurie, Danaé Metaxa +1
Large language models' (LLMs') outputs are shaped by opaque and frequently-changing company content moderation policies and practices. LLM moderation often takes the form of refusa…