Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data
arXiv:2508.09911 · doi:10.1145/3757707
Abstract
Data annotation underpins the success of modern AI, but the aggregation of crowd-collected datasets can harm the preservation of diverse perspectives in data. Difficult and ambiguous tasks cannot easily be collapsed into unitary labels. Prior work has shown that deliberation and discussion improve data quality and preserve diverse perspectives -- however, synchronous deliberation through crowdsourcing platforms is time-intensive and costly. In this work, we create a Socratic dialog system using Large Language Models (LLMs) to act as a deliberation partner in place of other crowdworkers. Against a benchmark of synchronous deliberation on two tasks (Sarcasm and Relation detection), our Socratic LLM encouraged participants to consider alternate annotation perspectives, update their labels as needed (with higher confidence), and resulted in higher annotation accuracy (for the Relation task where ground truth is available). Qualitative findings show that our agent's Socratic approach was effective at encouraging reasoned arguments from our participants, and that the intervention was well-received. Our methodology lays the groundwork for building scalable systems that preserve individual perspectives in generating more representative datasets.
To appear at CSCW 2025
References in corpus (17)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- Quality Control in Crowdsourcing: A Survey of Quality Attributes, Assessment Techniques and Assurance Actions
- Quizz: Targeted crowdsourcing with a billion (potential) users
- Jury Learning: Integrating Dissenting Voices into Machine Learning Models
- Like trainer, like bot? Inheritance of bias in algorithmic content moderation
- Deliberating with AI: Improving Decision-Making for the Future through Participatory AI Design and Stakeholder Deliberation
- Toward a Perspectivist Turn in Ground Truthing for Predictive Computing
- Writer-Defined AI Personas for On-Demand Feedback Generation
- Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
- Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement
- Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
- The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation
- Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty
- Evaluating Language Models for Generating and Judging Programming Feedback
- Learning From Crowdsourced Noisy Labels: A Signal Processing Perspective
- The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
- Instruct, Not Assist: LLM-based Multi-Turn Planning and Hierarchical Questioning for Socratic Code Debugging