565 citations · 879 across the 19 of their papers we have counts for
1 paper · 2 filters
Stephanie Lin, Jacob Hilton, Owain Evans
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including…