3 papers
cs.CL2025
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Jason Wei, Zhiqing Sun, Spencer Papay +7
We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently nav…
cs.CL2024
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung +5
We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First,…
cs.CL2024
GPT-4o System Card
OpenAI, :, Aaron Hurst +416
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…