works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.CL2026

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

Alicia Guerra, Yibo Hu

Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is…

cs.CL2026

Social Pressure Breaks Majority Voting in LLM Safety Panels

Yibo Hu, Jiaming Qu

Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but th…

cs.CR2026

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang +1

Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across sev…

cs.CR2026

When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems

Yibo Hu, Ren Wang

The paper shows that runtime monitors checking each step of multi-agent LLM systems can miss attacks that are split across agents, because each fragment looks benign on its own, an…

cs.CL2026

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

Yibo Hu, Jiaming Qu

LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even a…

cs.CL2026

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

Jiaming Qu, Lucheng Fu, Yibo Hu

Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answe…