1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Marina Mancoridis, Bec Weeks, Keyon Vafa +1
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated se…