1 paper
Marina Mancoridis, Bec Weeks, Keyon Vafa +1
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated se…