4 citations · 4 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Croissant: A Metadata Format for ML-Ready Datasets
Mubashara Akhtar, Omar Benjelloun, Costanza Conforti +28
Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that crea…
cs.LG2018
Efficient Augmentation via Data Subsampling
Michael Kuchnik, Virginia Smith
Data augmentation is commonly used to encode invariances in learning methods. However, this process is often performed in an inefficient manner, as artificial examples are created…