7 papers · 1 filter
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
Peiying Zhu, Sidi Chang
Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the…
When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
Peiying Zhu, Sidi Chang
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfa…
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Peiying Zhu, Sidi Chang
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable.…
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
Peiying Zhu, Sidi Chang
Outcome metrics can certify the wrong behavior. We study this failure in a two-hotel revenue-management simulator where Hotel A trains an agent against a fixed rule-based revenue-m…
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
Sidi Chang, Peiying Zhu, Yuxiao Chen
LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a delayed-ground-truth evaluation pro…
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
Peiying Zhu, Sidi Chang
Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral discipline. In hotel pricing with hidden compe…