2 papers
cs.LG2023
Comprehensive Benchmarking of Entropy and Margin Based Scoring Metrics for Data Selection
Anusha Sabbineni, Nikhil Anand, Maria Minakova
While data selection methods have been studied extensively in active learning, data pruning, and data augmentation settings, there is little evidence for the efficacy of these meth…
cs.LG2023
Influence Scores at Scale for Efficient Language Data Sampling
Nikhil Anand, Joshua Tan, Maria Minakova
Modern ML systems ingest data aggregated from diverse sources, such as synthetic, human-annotated, and live customer traffic. Understanding \textit{which} examples are important to…