4 papers
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Parsa Mazaheri, Kasra Mazaheri
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker re…
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Parsa Mazaheri
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a repres…
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
Parsa Mazaheri
One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the trajectory. We introduce RePoT (Recov…
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
Parsa Mazaheri, Kasra Mazaheri
Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but their evaluations often collapse behavior into final task…