From the 1 of 7 linked papers with an AI index.
7 papers
Tracing Agentic Failure from the Flow of Success
Samuel Yeh, Yiwen Zhu, Shaleen Deep +1
The paper introduces OAT, a lightweight unsupervised method that learns from successful LLM agent trajectories and detects error steps in failed runs by scoring deviations using ne…
Multi-Head Recurrent Memory Agents
Jiatong Li, Samuel Yeh, Sharon Li
Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their scalability, these agents exhibit…
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
Changdae Oh, Wendi Li, Seongheon Park +3
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irrever…
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
Seongheon Park, Wendi Li, Changdae Oh +4
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that…
DECOR: Auditing LLM Deception via Information Manipulation Theory
Linyue Cai, Samuel Yeh, Jwala Dhamala +2
Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detec…
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
Changdae Oh, Seongheon Park, To Eun Kim +8
Uncertainty quantification (UQ) for large language models (LLMs) is a key building block for safety guardrails of daily LLM applications. Yet, even as LLM agents are increasingly d…