6 papers
Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets
Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer +17
Croissant has emerged as the metadata standard for machine learning datasets, providing a structured, JSON-LD-based format that makes dataset discovery, automated ingestion, and re…
When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge
Andrés F. Castro Torres, Joan Giner-Miguelez, Mercè Crosas
The extent to which Artificial Intelligence (AI) technologies can trigger generalized paradigm shifts in science is unclear. Although these technologies have revolutionized data co…
The Software Diversity Card: A Framework for Reporting Diversity in Software Projects
Joan Giner-Miguelez, Sergio Morales, Sergio Cobos +3
Context: Interest in diversity in software development has significantly increased in recent years. Reporting on diversity in software projects can enhance user trust and assist re…
Enabling Content Management Systems as an Information Source in Model-driven Projects
Joan Giner-Miguelez, Abel Gómez, Jordi Cabot
Content Management Systems (CMSs) are the most popular tool when it comes to create and publish content across the web. Recently, CMSs have evolved, becoming \emph{headless}. Conte…
On the Readiness of Scientific Data for a Fair and Transparent Use in Machine Learning
Joan Giner-Miguelez, Abel Gómez, Jordi Cabot
To ensure the fairness and trustworthiness of machine learning (ML) systems, recent legislative initiatives and relevant research in the ML community have pointed out the need to d…
Croissant: A Metadata Format for ML-Ready Datasets
Mubashara Akhtar, Omar Benjelloun, Costanza Conforti +28
Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that crea…