collaborators

12 papers

cs.CR2026

When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems

Jialuo Chen, Lingqi Jiang, Xinhao Deng +7

Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a…

cs.AI2026

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

Jialuo Chen, Minghe Wang, Lingqi Jiang +7

LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workfl…

cs.AI2026

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Yunhao Feng, Ruixiao Lin, Ming Wen +12

LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed sa…

cs.CR2026

VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills

Ying Li, Yanju Chen, Hongbo Wen +5

Agentic systems increasingly act through third-party skills, allowing model-generated decisions to affect files, communication channels, and cyber-physical devices. These skills of…

cs.CR2026

Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies

Ruixiao Lin, Xinhao Deng, Qingming Li +12

Self-evolving LLM agent systems, which autonomously update their model parameters, memory, tools, and architectures, introduce a qualitatively new threat landscape in which adversa…

cs.LG2026

Constitutional On-Policy Safe Distillation

Ming Wen, Yuxuan Liu, Kun Yang +9

On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervis…