15 papers
AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming
Pingchuan Ma, Zhaoyu Wang, Zimo Ji +5
Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, their autonomy poses significant…
PathMark: Protecting Intellectual Property of Mixture-of-Expert LLMs via Path Watermarks
Yudong Gao, Qingyue Wang, Yuanyuan Yuan +4
Mixture-of-Experts (MoE) large language models represent high-value intellectual property, yet existing watermarking schemes designed for dense models fail on MoE architectures due…
Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware
Zimo Ji, Congying Xu, Zongjie Li +4
LLM coding agents increasingly rely on third-party agent skills from public marketplaces, which execute with the agent's privileges and create a software supply-chain attack surfac…
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
Xunguang Wang, Yuguang Zhou, Qingyue Wang +5
Large language models increasingly rely on explicit chain-of-thought reasoning to solve complex tasks, yet the safety of the reasoning process itself remains largely unaddressed. E…
EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
Yanbo Dai, Zhenlan Ji, Zongjie Li +1
Model editing techniques are essential for efficiently updating knowledge in large language models (LLMs). However, the effectiveness of existing approaches degrades in massive edi…
Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework
Zimo Ji, Daoyuan Wu, Wenyuan Jiang +5
Large Language Model (LLM)-based agent systems are increasingly deployed for complex real-world tasks but remain vulnerable to natural language-based attacks that exploit over-priv…