From the 1 of 10 linked papers with an AI index.
10 papers
The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity
Alicia Guerra, Yibo Hu
Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is…
Social Pressure Breaks Majority Voting in LLM Safety Panels
Yibo Hu, Jiaming Qu
Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but th…
Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang +1
Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across sev…
When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
Yibo Hu, Ren Wang
The paper shows that runtime monitors checking each step of multi-agent LLM systems can miss attacks that are split across agents, because each fragment looks benign on its own, an…
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
Yibo Hu, Jiaming Qu
LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even a…
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
Jiaming Qu, Lucheng Fu, Yibo Hu
Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answe…