3 papers
cs.CL2026
Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long Conversations
Chenhui Hu, Muhammed Salih, Sudipto Guha +1
Multi-turn jailbreaks can evade turn-level moderation by spreading unsafe intent across a dialogue through gradual escalation, reframing, and role manipulation. We address multi-tu…
cs.CL2026
Towards Atoms of Large Language Models
Chenhui Hu, Pengfei Cao, Yubo Chen +2
The fundamental representational units (FRUs) of large language models (LLMs) remain undefined, limiting further understanding of their underlying mechanisms. In this paper, we int…
cs.CL2025
Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models
Chenhui Hu, Pengfei Cao, Yubo Chen +2
Knowledge editing aims to update outdated or incorrect knowledge in large language models (LLMs). However, current knowledge editing methods have limited scalability for lifelong e…