2 papers
cs.CR2025
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
Zhuohang Long, Siyuan Wang, Shujun Liu +3
Jailbreak attacks, where harmful prompts bypass generative models' built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed,…
cs.CL2025
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
Shengbin Yue, Ting Huang, Zheng Jia +5
Large Language Models (LLMs) have significantly advanced legal intelligence, but the scarcity of scenario data impedes the progress toward interactive legal scenarios. This paper i…