Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
PCA-guided Activation Scaling for Monotonic Bidirectional Control over LLM Sycophancy
Zheng Chen, Zhaoxin Feng, Yip Tin Po +3
Large language models (LLMs) exhibit sycophancy, a tendency to agree with user beliefs regardless of factual accuracy. This can reinforce misconceptions, but eliminating it entirel…
cs.CL2026
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
Zhaoxin Feng, Zheng Chen, Jianfei Ma +3
Alignment techniques often inadvertently induce sycophancy in LLMs. While prior studies studied this behaviour in direct-answer settings, the role of Chain-of-Thought (CoT) reasoni…