2 papers
cs.LG2026
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
Adarsh Kumarappan, Ananya Mujoo
LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced s…
cs.LG2026
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
Adarsh Kumarappan, Ananya Mujoo
Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to…