1 citations · 1 across the 1 of their papers we have counts for
1 paper
Lukas Haas, Gal Yona, Giovanni D'Antonio +2
We introduce SimpleQA Verified, a 1,000-prompt benchmark for evaluating Large Language Model (LLM) short-form factuality based on OpenAI's SimpleQA. It addresses critical limitatio…