2 papers
cs.CY2026
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
Marina Mancoridis, Zoë Hitzig
Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pip…
cs.CL2025
Potemkin Understanding in Large Language Models
Marina Mancoridis, Bec Weeks, Keyon Vafa +1
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated se…