3 papers
cs.LG2025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy
Rudrajit Das, Inderjit S. Dhillon, Alessandro Epasto +5
The performance of a model trained with noisy labels is often improved by simply \textit{retraining} the model with its \textit{own predicted hard labels} (i.e., 1/0 labels). Yet,…
cs.LG2025
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
Anish Acharya, Sujay Sanghavi, Alexandros G. Dimakis +1
Data pruning -- the combinatorial task of selecting a small and representative subset from a large dataset, is crucial for mitigating the enormous computational costs associated wi…
cs.LG2025
Geometric Median (GM) Matching for Robust Data Pruning
Anish Acharya, Inderjit S Dhillon, Sujay Sanghavi
Large-scale data collections in the wild, are invariably noisy. Thus developing data pruning strategies that remain robust even in the presence of corruption is critical in practic…