From the 1 of 28 linked papers with an AI index.
28 papers
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
Hongwei Yao, Yiming Liu, Meihui Chen +6
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBe…
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
Weiwei Qi, Zefeng Wu, Zhilin Guo +5
Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…
DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
Zefeng Wu, Weiwei Qi, Jielong Chen +6
The paper introduces DataShield, a framework that detects risky fine‑tuning data for large language models by aligning safety‑critical semantic subspaces across multiple safety‑ali…
NonTextual Target Attack
Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint
Shuo Shao, Haozhe Zhu, Yiming Li +3
Model fingerprinting has emerged as a crucial mechanism for safeguarding the intellectual property of open-source models, offering a non-intrusive approach that requires no modific…
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
Churui Zeng, Weiwei Qi, Kedong Xiu +5
The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…