4 papers
Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations
Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon +8
Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and bri…
CUBE: A Standard for Unifying Agent Benchmarks
Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Nicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold +7
We propose a new algorithm for fine-tuning large language models using reinforcement learning. Tapered Off-Policy REINFORCE (TOPR) uses an asymmetric, tapered variant of importance…
Mitigating Downstream Model Risks via Model Provenance
Keyu Wang, Abdullah Norozi Iranzad, Scott Schaffter +3
Research and industry are rapidly advancing the innovation and adoption of foundation model-based systems, yet the tools for managing these models have not kept pace. Understanding…