Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
Xiaoyu Wen, Zhida He, Han Qi +7
Ensuring robust safety alignment is crucial for Large Language Models (LLMs), yet existing defenses often lag behind evolving adversarial attacks due to their \textbf{reliance on s…
cs.AI2026
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
Xingcheng Xu, Jingjing Qu, Qiaosheng Zhang +4
The rapid deployment of Large Language Models and AI agents across critical societal and technical domains is hindered by persistent behavioral pathologies including sycophancy, ha…