5 papers
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker
End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure…
Belief Engine: Configurable and Inspectable Stance Dynamics in Multi-Agent LLM Deliberation
Joshua C. Yang, Maurice Flechtner, Damian Dailisan +1
LLM-based agents are increasingly used to simulate deliberative interactions such as negotiation, conflict resolution, and multi-turn opinion exchange. Yet generated transcripts of…
AI Fact-Checking in the Wild: A Field Evaluation of LLM-Written Community Notes on X
Haiwen Li, Michiel A. Bakker
Large language models (LLMs) show promising capabilities for fact-checking, yet prior work evaluates them only in controlled offline settings using benchmarks or crowdworker judgme…
Community Moderation and the New Epistemology of Fact Checking on Social Media
Isabelle Augenstein, Michiel Bakker, Tanmoy Chakraborty +13
Social media platforms have traditionally relied on internal moderation teams and partnerships with independent fact-checking organizations to identify and flag misleading content.…
Tell Me Why: Incentivizing Explanations
Siddarth Srinivasan, Ezra Karger, Michiel Bakker +1
Common sense suggests that when individuals explain why they believe something, we can arrive at more accurate conclusions than when they simply state what they believe. Yet, there…