Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
Yutong Zhang, Jianshuo Dong, Peng Xu +5
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harm…
cs.CL2025
Towards Understanding the Cognitive Habits of Large Reasoning Models
Jianshuo Dong, Yujia Fu, Chuanrui Hu +2
Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monito…