1 citations · 1 across the 1 of their papers we have counts for
1 paper
Marina Mancoridis, Bec Weeks, Keyon Vafa +1
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated se…