2 papers
cs.AI2026
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Lanbo Lin, Jiayao Liu, Tianyuan Yang +5
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail…
cs.SE2025
Automated Statistical Testing and Certification of a Reliable Model-Coupling Server for Scientific Computing
Seth Wolfgang, Lan Lin, Fengguang Song
Sequence-based specification and usage-driven statistical testing are designed for rigorous and cost-effective software development, offering a semi-formal approach to assessing th…