collaborators

11 papers

cs.LG2026

Safin-1: Safety from Within through Memory-Native State Evolution

Ming Zhang, Kaisen Yang, Shu Yu +15

Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic proper…

cs.CR2026

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

Zhida He, Xiaoyu Wen, Han Qi +5

Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on att…

cs.AI2026

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

Xiaoyu Wen, Jiajia Li, Zhida He +11

Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematica…

cs.CL2026

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

Zhida He, Xia Hu, Baichen Le +20

Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the…

cs.AI2026

Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

Zhida He, Xiaoyu Wen, Han Qi +7

Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based multi-turn jailbreak method…

cs.CR2026

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

Chang Jin, An Wang, Zeming Wei +7

Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environm…