2 papers
cs.CR2026
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
Xuan Luo, Yue Wang, Geng Tu +2
In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the mod…
cs.CR2026
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
Xuan Luo, Yue Wang, Zefeng He +3
This study reveals a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses…