From the 1 of 6 linked papers with an AI index.
6 papers
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
Sidi Chang, Peiying Zhu, Yuxiao Chen +1
The paper investigates how the wording of evaluation rubrics and the choice of metrics affect the reliability of supervised financial NLP benchmarks, using a Japanese implicit‑comm…
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
Peiying Zhu, Sidi Chang
Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral discipline. In hotel pricing with hidden compe…
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Peiying Zhu, Sidi Chang
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable.…
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
Peiying Zhu, Sidi Chang
Outcome metrics can certify the wrong behavior. We study this failure in a two-hotel revenue-management simulator where Hotel A trains an agent against a fixed rule-based revenue-m…
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
Sidi Chang, Peiying Zhu, Yuxiao Chen
LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a delayed-ground-truth evaluation pro…
The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries
Peiying Zhu, Sidi Chang
When a traveler asks an AI search engine to recommend a hotel, which sources get cited -- and does query framing matter? We audit 1,357 grounding citations from Google Gemini acros…