4 papers
When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
Peiying Zhu, Sidi Chang
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfa…
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Peiying Zhu, Sidi Chang
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable.…
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
Sidi Chang, Peiying Zhu, Yuxiao Chen
LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a delayed-ground-truth evaluation pro…
The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries
Peiying Zhu, Sidi Chang
When a traveler asks an AI search engine to recommend a hotel, which sources get cited -- and does query framing matter? We audit 1,357 grounding citations from Google Gemini acros…