agent evaluation 1benchmark consolidation 1capability scaling 1cross-benchmark analysis 1dataset infrastructure 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Predicting Task Difficulty Without Rollouts
Stefan Krsteski, Charlotte Meyer
Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simula…
cs.AI2026
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2
The paper introduces Messier, a unified corpus of 957,253 standardized evaluation records spanning thousands of agents, tasks, and benchmarks, to enable cross‑benchmark analysis an…