3 papers
cs.CV2024
Label Errors in the Tobacco3482 Dataset
Gordon Lim, Stefan Larson, Kevin Leach
Tobacco3482 is a widely used document classification benchmark dataset. However, our manual inspection of the entire dataset uncovers widespread ontological issues, especially larg…
cs.HC2024
Towards Fair Pay and Equal Work: Imposing View Time Limits in Crowdsourced Image Classification
Gordon Lim, Stefan Larson, Yu Huang +1
Crowdsourcing is a common approach to rapidly annotate large volumes of data in machine learning applications. Typically, crowd workers are compensated with a flat rate based on an…
cs.LG2024
Robust Testing for Deep Learning using Human Label Noise
Gordon Lim, Stefan Larson, Kevin Leach
In deep learning (DL) systems, label noise in training datasets often degrades model performance, as models may learn incorrect patterns from mislabeled data. The area of Learning…