4 papers
Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
Yuxuan Gao, Megan Wang, Yi Ling Yu
We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guarantees for forecasted quality…
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
Yuxuan Gao, Megan Wang, Yi Ling Yu +2
We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau-bench, BFCL multi-turn), a…
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
Yuxuan Gao, Megan Wang, Yi Ling Yu
Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model,…
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
Yuxuan Gao, Megan Wang, Yi Ling Yu
Static benchmarks measure what AI agents can do at a fixed point in time but not how they are adopted, maintained, or experienced in deployment. We introduce AgentPulse, a continuo…