50 citations · 237 across the 42 of their papers we have counts for
26 papers · 1 filter
MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection
Yue Wang, Yi Liu, Gelei Deng +4
Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This creates a direct distribution channel fo…
Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks
Xiaoyan Feng, Yanjun Zhang, He Zhang +2
Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adver…
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics
Hangtao Zhang, Yucheng Zhao, Sishun Liu +8
Jailbreak prompts can bypass alignment guardrails in large language models (LLMs) and elicit unsafe outputs, making reliable deployment-time detection critical. Prior detection app…
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
Yubin Qu, Yi Liu, Gelei Deng +4
A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while the task still completes. We cal…
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
Ruoqi Guo, Yi Liu, Gelei Deng +7
Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot re…
How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
Zhihao Chen, Ying Zhang, Yi Liu +7
Large Language Model (LLM) agents increasingly rely on third-party skills that operate within privileged execution environments and routinely handle sensitive credentials, yet how…