4 papers
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Parsa Mazaheri
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a repres…
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
Parsa Mazaheri
One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the trajectory. We introduce RePoT (Recov…
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
Parsa Mazaheri, Kasra Mazaheri
Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but their evaluations often collapse behavior into final task…
Don't Believe Everything You Read: Enhancing Summarization Interpretability through Automatic Identification of Hallucinations in Large Language Models
Priyesh Vakharia, Devavrat Joshi, Meenal Chavan +3
Large Language Models (LLMs) are adept at text manipulation -- tasks such as machine translation and text summarization. However, these models can also be prone to hallucination, w…