activity
20242026
collaborators

10 papers

cs.CR2026

SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents

Xiaojun Jia, Jie Liao, Simeng Qin +5

Agent skills extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources, improving reusability but creating a new supply-chain attack surface. A…

cs.CV2026

Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models

Songlong Xing, Weijie Wang, Zhengyu Zhao +3

Despite their impressive zero-shot abilities, vision-language models such as CLIP have been shown to be susceptible to adversarial attacks. To enhance its adversarial robustness, r…

cs.CR2026

LLM Jailbreak Detection for (Almost) Free!

Guorui Chen, Yifan Xia, Xiaojun Jia +3

Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak…

cs.CL2026

Can Editing LLMs Inject Harm?

Canyu Chen, Baixiang Huang, Zekun Li +12

Large Language Models (LLMs) have emerged as a new information channel. Meanwhile, one critical but under-explored question is: Is it possible to bypass the safety alignment and in…

cs.AI2025

Reimagining Safety Alignment with An Image

Yifan Xia, Guorui Chen, Wenqian Yu +3

Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusal of benign queries due to ri…

cs.AI2025

Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?

Junchi Yu, Yujie Liu, Jindong Gu +2

Retrieval-Augmented Generation (RAG) based on knowledge graphs (KGs) enhances large language models (LLMs) by providing structured and interpretable external knowledge. However, ex…