1 citations · 1 across the 15 of their papers we have counts for
9 papers · 1 filter
Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits
Yanhang Li, Zhichao Fan, Zexin Zhuang
Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence. We argu…
Do Time Series Foundation Model Benchmarks Hide Regime-Dependent Failures? Evidence from Traffic Speed Forecasting
Yingshuo Wang, Xian Sun, Lingdong Kong +4
Standard benchmarks evaluate time series foundation models (TSFMs) using aggregate metrics, but these can mask severe failures in critical operating regimes. We introduce regime-st…
Embedding Foundation Model Predictions in Discrete-Choice Models with Structural Guarantees
Yingshuo Wang, Xian Sun, Yanhang Li +2
Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks require: raising a price can increa…
When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents
Yanhang Li, Zhichao Fan, Zexin Zhuang
Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for flagging indirect prompt in…
Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation
Xian Sun, Wei Gao, Yingshuo Wang +9
Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support…
Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice
Yingshuo Wang, Xian Sun, Yanhang Li +2
Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks require: raising a price sometimes…