Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness
Sawyer Zhang, Alexander Wang, Sophie Lei
End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present layer-isolated evaluation: a d…
cs.CL2026
Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents
Sawyer Zhang, Alexander Wang, Sophie Lei
LLM-as-judge is the default instrument for evaluating conversational agents, yet its reliability is almost always reported as agreement with human ratings, not recall of real defec…