activity
20242026
collaborators

6 papers

cs.CR2026

Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills

Lijia Lv, Xuehai Tang, Jie Wen +2

Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering int…

cs.CL2026

FABLE: Fine-grained Fact Anchoring for Unstructured Model Editing

Peng Wang, Biyu Zhou, Xuehai Tang +2

Unstructured model editing aims to update models with real-world text, yet existing methods often memorize text holistically without reliable fine-grained fact access. To address t…

cs.CL2025

Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs

Xikang Yang, Biyu Zhou, Xuehai Tang +2

Large Language Models (LLMs) demonstrate impressive capabilities across a wide range of tasks, yet their safety mechanisms remain susceptible to adversarial attacks that exploit co…

cs.CL2025

LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing

Peng Wang, Biyu Zhou, Xuehai Tang +2

Large Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates. However, current mainstream locat…

cs.CL2025

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers

Liang Lin, Zhihao Xu, Xuehai Tang +5

The safety of large language models (LLMs) has garnered significant research attention. In this paper, we argue that previous empirical studies demonstrate LLMs exhibit a propensit…

cs.LG2024

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Xikang Yang, Xuehai Tang, Jizhong Han +1

The widespread deployment of large language models (LLMs) across various domains has showcased their immense potential while exposing significant safety vulnerabilities. A major co…