Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Potemkin Understanding in Large Language Models
Marina Mancoridis, Bec Weeks, Keyon Vafa +1
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated se…
cs.CL2024
Evaluating the World Model Implicit in a Generative Model
Keyon Vafa, Justin Y. Chen, Ashesh Rambachan +2
Recent work suggests that large language models may implicitly learn world models. How should we assess this possibility? We formalize this question for the case where the underlyi…
cs.CL2024
Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function
Keyon Vafa, Ashesh Rambachan, Sendhil Mullainathan
What makes large language models (LLMs) impressive is also what makes them hard to evaluate: their diversity of uses. To evaluate these models, we must understand the purposes they…