Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations
Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon +8
Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and bri…
cs.AI2026
CUBE: A Standard for Unifying Agent Benchmarks
Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…