6 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.CL2025★ 1 cited
Potemkin Understanding in Large Language Models
Marina Mancoridis, Bec Weeks, Keyon Vafa +1
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated se…
cs.CL2024★ 6 cited
Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function
Keyon Vafa, Ashesh Rambachan, Sendhil Mullainathan
What makes large language models (LLMs) impressive is also what makes them hard to evaluate: their diversity of uses. To evaluate these models, we must understand the purposes they…
cs.CL2023★ 1 cited
An Invariant Learning Characterization of Controlled Text Generation
Carolina Zheng, Claudia Shi, Keyon Vafa +2
Controlled generation refers to the problem of creating text that contains stylistic or semantic attributes of interest. Many approaches reduce this problem to training a predictor…