1 paper · 1 filter
Yutong Zhang, Jianshuo Dong, Peng Xu +5
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harm…