3 papers
cs.LG2026
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Laksh Advani
LLM agents can fail silently by asserting task completion when the environment state shows otherwise. We study this failure mode, false success, across two agent benchmarks: 9,876…
cs.LG2026
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
Laksh Advani
Autonomous LLM agents generate multi-step action plans that can fail due to contextual misalignment or structural incoherence. Existing anomaly detection methods are ill-suited for…
cs.LG2026
When Small Models Are Right for Wrong Reasons: Process Verification for Trustworthy Agents
Laksh Advani
Deploying small language models (7-9B parameters) as autonomous agents requires trust in their reasoning, not just their outputs. We reveal a critical reliability crisis: 50-69\% o…