papers

Publications (21)

cs.AI2026

Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon +8

Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and bri…

cs.LG2024

Generative Fractional Diffusion Models

Gabriel Nobis, Maximilian Springenberg, Marco Aversa +11

We introduce the first continuous-time score-based generative model that leverages fractional diffusion processes for its underlying dynamics. Although diffusion models have excell…

cs.LG2023

Data Models for Dataset Drift Controls in Machine Learning With Optical Images

Luis Oala, Marco Aversa, Gabriel Nobis +10

Camera images are ubiquitous in machine learning research. They also play a central role in the delivery of important services spanning medicine and environmental surveying. Howeve…

cs.CY2025

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Shaona Ghosh, Heather Frase, Adina Williams +99

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…

cs.CV2026

Autoguided Online Data Curation for Diffusion Model Training

Valeria Pais, Luis Oala, Daniele Faccio +1

The costs of generative model compute rekindled promises and hopes for efficient data curation. In this work, we investigate whether recently developed autoguidance and online data…

cs.IR2024

A Standardized Machine-readable Dataset Documentation Format for Responsible AI

Nitisha Jain, Mubashara Akhtar, Joan Giner-Miguelez +15

Data is critical to advancing AI technologies, yet its quality and documentation remain significant challenges, leading to adverse downstream effects (e.g., potential biases) in AI…