26 citations · 38 across the 13 of their papers we have counts for
8 papers · 1 filter
The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity
Alicia Guerra, Yibo Hu
Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is…
Social Pressure Breaks Majority Voting in LLM Safety Panels
Yibo Hu, Jiaming Qu
Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but th…
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
Yibo Hu, Jiaming Qu
LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even a…
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
Jiaming Qu, Lucheng Fu, Yibo Hu
Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answe…
Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA
Maroof Kousar, Yibo Hu
Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the-counter (OTC) medication. Yet…
When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding
Zixian He, Bharath Raahul Murugesan, Patrick Brandt +1
High accuracy does not necessarily make an LLM a faithful coder. This issue matters because many social-science studies rely on expert-written codebooks to turn text into structure…