2 papers
cs.AI2026
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2
Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts…
cs.LG2026
Predicting Task Difficulty Without Rollouts
Stefan Krsteski, Charlotte Meyer
Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simula…