Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
Sagar Chethan Kumar, Rohith Kanathur, Dhaval Patel +1
Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks…
cs.AI2026
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…