1 paper · 1 filter
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski +1
Pass@k is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especially when the number of trials (sampl…