4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.IR2024★ 4 cited
A Standardized Machine-readable Dataset Documentation Format for Responsible AI
Nitisha Jain, Mubashara Akhtar, Joan Giner-Miguelez +15
Data is critical to advancing AI technologies, yet its quality and documentation remain significant challenges, leading to adverse downstream effects (e.g., potential biases) in AI…
cs.LG2024
Croissant: A Metadata Format for ML-Ready Datasets
Mubashara Akhtar, Omar Benjelloun, Costanza Conforti +28
Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that crea…