collaborators

12 papers

cs.SE2026

Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures

Zihao Zheng, Baichuan Li, Junyi Yao +1

Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic ben…

cs.CL2026

Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

Junyi Yao, Baichuan Li, Zihao Zheng +1

Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet m…

cs.CL2026

Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction

Zihao Zheng, Baichuan Li, Junyi Yao +1

Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indi…

cs.CL2026

Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

Baichuan Li, Junyi Yao, Zihao Zheng

Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interactio…

cs.AI2026

Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems

Junyi Yao, Zihao Zheng

Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary i…

cs.AI2026

Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

Junyi Yao, Zihao Zheng, Baichuan Li

Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent c…