6 papers
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
Kaiyue Yang, Yuyan Bu, Jingwei Yi +5
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on saf…
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
Lijia Lv, Xuehai Tang, Jie Wen +2
Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering int…
FABLE: Fine-grained Fact Anchoring for Unstructured Model Editing
Peng Wang, Biyu Zhou, Xuehai Tang +2
Unstructured model editing aims to update models with real-world text, yet existing methods often memorize text holistically without reliable fine-grained fact access. To address t…
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
Xikang Yang, Biyu Zhou, Xuehai Tang +2
Large Language Models (LLMs) demonstrate impressive capabilities across a wide range of tasks, yet their safety mechanisms remain susceptible to adversarial attacks that exploit co…
LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing
Peng Wang, Biyu Zhou, Xuehai Tang +2
Large Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates. However, current mainstream locat…
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
Liang Lin, Zhihao Xu, Xuehai Tang +5
The safety of large language models (LLMs) has garnered significant research attention. In this paper, we argue that previous empirical studies demonstrate LLMs exhibit a propensit…