works on

From the 1 of 28 linked papers with an AI index.

activity
20242026
collaborators

28 papers

cs.CR2026

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Hongwei Yao, Yiming Liu, Meihui Chen +6

Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBe…

cs.CR2026

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Weiwei Qi, Zefeng Wu, Zhilin Guo +5

Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…

cs.CR2026

DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment

Zefeng Wu, Weiwei Qi, Jielong Chen +6

The paper introduces DataShield, a framework that detects risky fine‑tuning data for large language models by aligning safety‑critical semantic subspaces across multiple safety‑ali…

cs.CR2026

NonTextual Target Attack

Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6

Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…

cs.CR2026

FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint

Shuo Shao, Haozhe Zhu, Yiming Li +3

Model fingerprinting has emerged as a crucial mechanism for safeguarding the intellectual property of open-source models, offering a non-intrusive approach that requires no modific…

cs.CR2026

TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking

Churui Zeng, Weiwei Qi, Kedong Xiu +5

The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…