3 papers
cs.AI2026
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
cs.CL2026
Failure of contextual invariance in large language models
Sagar Kumar, Ariel Flint, Luca Maria Aiello +1
Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses. Here, we test this assumpti…
cs.AI2026
Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
Sagar Chethan Kumar, Rohith Kanathur, Dhaval Patel +1
Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks…