4 papers
Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures
Zihao Zheng, Baichuan Li, Junyi Yao +1
Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic ben…
Toward Workflow-Aware Benchmarking for Healthcare NLP Agents
Junyi Yao, Baichuan Li, Zihao Zheng +1
Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet m…
Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction
Zihao Zheng, Baichuan Li, Junyi Yao +1
Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indi…
Perturbation Sensitivity of Maximum-Likelihood Pairwise Ranking in Computational Decision Systems
Junyi Yao, Zihao Zheng, Jiayu Long
Maximum-likelihood pairwise ranking is a com- mon computational mechanism for prioritization, reputation estimation, and comparison-driven decision support. Despite its broad use,…