6 citations · 8 across the 2 of their papers we have counts for
4 papers
Identifying Incorrect Annotations in Multi-Label Classification Data
Aditya Thyagarajan, Elías Snorrason, Curtis Northcutt +1
In multi-label classification, each example in a dataset may be annotated as belonging to one or more classes (or none of the classes). Example applications include image (or docum…
Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
Curtis G. Northcutt, Anish Athalye, Jonas Mueller
We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these labe…
Rapformer: Conditional Rap Lyrics Generation with Denoising Autoencoders
Nikola I. Nikolov, Eric Malmi, Curtis G. Northcutt +1
The ability to combine symbols to generate language is a defining characteristic of human intelligence, particularly in the context of artistic story-telling through lyrics. We dev…
Comment Ranking Diversification in Forum Discussions
Curtis G. Northcutt, Kimberly A. Leon, Naichun Chen
Viewing consumption of discussion forums with hundreds or more comments depends on ranking because most users only view top-ranked comments. When comments are ranked by an ordered…