3 papers
cs.HC2026
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Ro Encarnación, Tina Behzad, Emma Lurie +1
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a sing…
cs.CY2026
The Beginning of ChatGPT Ads
Emma Lurie, Ro Encarnación, Sorelle A. Friedler +1
This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examin…
cs.AI2026
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian +1
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms…