7 papers
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
Yunhao Feng, Ruixiao Lin, Ming Wen +12
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed sa…
Constitutional On-Policy Safe Distillation
Ming Wen, Yuxuan Liu, Kun Yang +9
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervis…
Orthogonal Concept Erasure for Diffusion Models
Yuhao Sun, Lingyun Yu, Haoxiang Xu +3
Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face significant limitations. While trai…
UpSafeC: Upcycling for Controllable Safety in Large Language Models
Yuhao Sun, Zhuoer Xu, Shiwen Cui +4
Large Language Models (LLMs) have achieved remarkable progress across a wide range of tasks, but remain vulnerable to safety risks such as harmful content generation and jailbreak…
TraceLLM: Security Diagnosis Through Traces and Smart Contracts in Ethereum
Shuzheng Wang, Yue Huang, Zhuoer Xu +2
Ethereum smart contracts hold tens of billions of USD in DeFi and NFTs, yet comprehensive security analysis remains difficult due to unverified code, proxy-based architectures, and…
Agent Safety Alignment via Reinforcement Learning
Zeyang Sha, Hanling Tian, Zhuoer Xu +3
The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents,…