1 paper · 1 filter
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2
The paper introduces Messier, a unified corpus of 957,253 standardized evaluation records spanning thousands of agents, tasks, and benchmarks, to enable cross‑benchmark analysis an…