1 paper
Riccardo Fogliato, Pratik Patil, Nil-Jana Akpinar +1
How can we precisely estimate a large language model's (LLM) accuracy on questions belonging to a specific topic within a larger question-answering dataset? The standard direct est…