Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
arXiv:2302.02990 · doi:10.1145/3544548.3581431
Abstract
Human data labeling is an important and expensive task at the heart of supervised learning systems. Hierarchies help humans understand and organize concepts. We ask whether and how concept hierarchies can inform the design of annotation interfaces to improve labeling quality and efficiency. We study this question through annotation of vaccine misinformation, where the labeling task is difficult and highly subjective. We investigate 6 user interface designs for crowdsourcing hierarchical labels by collecting over 18,000 individual annotations. Under a fixed budget, integrating hierarchies into the design improves crowdsource workers' F1 scores. We attribute this to (1) Grouping similar concepts, improving F1 scores by +0.16 over random groupings, (2) Strong relative performance on high-difficulty examples (relative F1 score difference of +0.40), and (3) Filtering out obvious negatives, increasing precision by +0.07. Ultimately, labeling schemes integrating the hierarchy outperform those that do not - achieving mean F1 of 0.70.
To appear in CHI-2023
References in corpus (8)
- The Kinetics Human Action Video Dataset
- Galaxy Zoo : Morphologies derived from visual inspection of galaxies from the Sloan Digital Sky Survey
- LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
- FMA: A Dataset For Music Analysis
- Embracing Error to Enable Rapid Crowdsourcing
- A Permutation-based Model for Crowd Labeling: Optimal Estimation and Robustness
- Distributed NLI: Learning to Predict Human Opinion Distributions for Language Reasoning
- Much Ado About Time: Exhaustive Annotation of Temporal Data