From the 2 of 22 linked papers with an AI index.
22 papers
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
Xiaoyu Wen, Jiajia Li, Zhida He +11
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematica…
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
Xinzhe Huang, Biwu Yao, Kedong Xiu +4
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
Zhengze Huang, Luyang Yu, Di Hong +5
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
Yuxuan Huang, Xingyu Zeng, Tianhang Zheng +1
Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradig…
AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
Ruoyu Wang, Heng Zhao, Renjie Wu +4
The paper presents AgentSnare, a system that dynamically creates deceptive decoy environments to mislead and delay autonomous penetration testing agents powered by large language m…
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
Weiwei Qi, Zefeng Wu, Zhilin Guo +5
Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…