A Permutation-based Model for Crowd Labeling: Optimal Estimation and Robustness
arXiv:1606.09632 · doi:10.1109/TIT.2020.3045613
Abstract
The task of aggregating and denoising crowd-labeled data has gained increased significance with the advent of crowdsourcing platforms and massive datasets. We propose a permutation-based model for crowd labeled data that is a significant generalization of the classical Dawid-Skene model, and introduce a new error metric by which to compare different estimators. We derive global minimax rates for the permutation-based model that are sharp up to logarithmic factors, and match the minimax lower bounds derived under the simpler Dawid-Skene model. We then design two computationally-efficient estimators: the WAN estimator for the setting where the ordering of workers in terms of their abilities is approximately known, and the OBI-WAN estimator where that is not known. For each of these estimators, we provide non-asymptotic bounds on their performance. We conduct synthetic simulations and experiments on real-world crowdsourcing data, and the experimental results corroborate our theoretical findings.
in IEEE Transactions on Information Theory (online), 2020
References in corpus (8)
- Concentration around the mean for maxima of empirical processes
- Regularized Minimax Conditional Entropy for Crowdsourcing
- Exact Exponent in Optimal Rates for Crowdsourcing
- PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review
- Estimating the Accuracies of Multiple Classifiers Without Labeled Data
- Your 2 is My 1, Your 3 is My 9: Handling Arbitrary Miscalibrations in Ratings
- On Testing for Biases in Peer Review
- Breaking the Barrier: Faster Rates for Permutation-based Models in Polynomial Time
Cited by in corpus (13)
- Adversarial Crowdsourcing Through Robust Rank-One Matrix Completion
- Worst-case vs Average-case Design for Estimation from Fixed Pairwise Comparisons
- Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
- Reducing Crowdsourcing to Graphon Estimation, Statistically
- Breaking the Barrier: Faster Rates for Permutation-based Models in Polynomial Time
- Universal Clustering via Crowdsourcing
- Iterative Bayesian Learning for Crowdsourced Regression
- A Worker-Task Specialization Model for Crowdsourcing: Efficient Inference and Fundamental Limits
- CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription
- Optimal detection of the feature matching map in presence of noise and outliers
- Crowdsourced Labeling for Worker-Task Specialization Model
- Towards Optimal Estimation of Bivariate Isotonic Matrices with Unknown Permutations
- Optimal rates for ranking a permuted isotonic matrix in polynomial time