1 paper
Jason Wei, Nguyen Karina, Hyung Won Chung +5
We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First,…