Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking
arXiv:2504.19940 · doi:10.1016/j.osnem.2025.100326
Abstract
The growing spread of online misinformation has created an urgent need for scalable, reliable fact-checking solutions. Crowdsourced fact-checking - where non-experts evaluate claim veracity - offers a cost-effective alternative to expert verification, despite concerns about variability in quality and bias. Encouraged by promising results in certain contexts, major platforms such as X (formerly Twitter), Facebook, and Instagram have begun shifting from centralized moderation to decentralized, crowd-based approaches. In parallel, advances in Large Language Models (LLMs) have shown strong performance across core fact-checking tasks, including claim detection and evidence evaluation. However, their potential role in crowdsourced workflows remains unexplored. This paper investigates whether LLM-powered generative agents - autonomous entities that emulate human behavior and decision-making - can meaningfully contribute to fact-checking tasks traditionally reserved for human crowds. Using the protocol of La Barbera et al. (2024), we simulate crowds of generative agents with diverse demographic and ideological profiles. Agents retrieve evidence, assess claims along multiple quality dimensions, and issue final veracity judgments. Our results show that agent crowds outperform human crowds in truthfulness classification, exhibit higher internal consistency, and show reduced susceptibility to social and cognitive biases. Compared to humans, agents rely more systematically on informative criteria such as Accuracy, Precision, and Informativeness, suggesting a more structured decision-making process. Overall, our findings highlight the potential of generative agents as scalable, consistent, and less biased contributors to crowd-based fact-checking systems.
This paper has been published in Online Social Networks and Media (https://doi.org/10.1016/j.osnem.2025.100326). Please cite the published version accordingly
References in corpus (16)
- Survey of Hallucination in Natural Language Generation
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Crowdsourced Fact-Checking at Twitter: How Does the Crowd Compare With Experts?
- The Many Dimensions of Truthfulness: Crowdsourcing Misinformation Assessments on a Multidimensional Scale
- Aligning Large Language Models with Human: A Survey
- Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation
- Customized large language models can outperform Community Notes in correcting misinformation
- Large Language Model Agent for Fake News Detection
- FACTors: A New Dataset for Studying the Fact-checking Ecosystem
- Generative Large Language Models in Automated Fact-Checking: A Survey
- Reinforcement Retrieval Leveraging Fine-grained Feedback for Fact Checking News Claims with Black-Box LLM
- Moral Alignment for LLM Agents
- Evidence-based Interpretable Open-domain Fact-checking with Large Language Models
- Towards Automated Fact-Checking of Real-World Claims: Exploring Task Formulation and Assessment with LLMs
- The Real, the Better: Aligning Large Language Models with Online Human Behaviors