3 papers
cs.CL2025
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Jason Wei, Zhiqing Sun, Spencer Papay +7
We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently nav…
cs.CL2025
Deliberative Alignment: Reasoning Enables Safer Language Models
Melody Y. Guan, Manas Joglekar, Eric Wallace +12
As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introdu…
cs.CL2024
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung +5
We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First,…