works on

From the 2 of 22 linked papers with an AI index.

collaborators

22 papers

cs.AI2026

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

Xiaoyu Wen, Jiajia Li, Zhida He +11

Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematica…

cs.LG2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

Xinzhe Huang, Biwu Yao, Kedong Xiu +4

Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…

cs.AI2026

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

Zhengze Huang, Luyang Yu, Di Hong +5

Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…

cs.CR2026

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Yuxuan Huang, Xingyu Zeng, Tianhang Zheng +1

Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradig…

cs.CR2026

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

Ruoyu Wang, Heng Zhao, Renjie Wu +4

The paper presents AgentSnare, a system that dynamically creates deceptive decoy environments to mislead and delay autonomous penetration testing agents powered by large language m…

cs.CR2026

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Weiwei Qi, Zefeng Wu, Zhilin Guo +5

Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…