collaborators

9 papers

cs.CL2026

Social Pressure Breaks Majority Voting in LLM Safety Panels

Yibo Hu, Jiaming Qu

Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but th…

cs.CL2026

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

Yibo Hu, Jiaming Qu

LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even a…

cs.CL2026

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text

Hongbo Du, Zixin Lu, Jiaming Qu

Large language models (LLMs) are increasingly used for clinical text tasks such as summarization and revision. While most studies evaluate the fluency and coherence of LLM-generate…

cs.CL2026

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

Jiaming Qu, Lucheng Fu, Yibo Hu

Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answe…

cs.LG2026

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

Chen Ling, Pei Chen, Albert Guan +4

Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other components of the agent pipeline.…

cs.CL2026

Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues

Jiaming Qu, Mengtian Guo, Yue Wang

Deceptive reviews mislead consumers, harm businesses, and undermine trust in online marketplaces. Machine learning classifiers can learn from large amounts of data to distinguish d…