4 papers
PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
Junhui Wang, Hangtao Zhang, Zhirun Zheng +5
Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are exp…
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
Tianxin Chen, Wenbo Jiang, Hongqiao Chen +2
Text-to-image (T2I) diffusion models are widely adopted for their strong generative capabilities, yet remain vulnerable to backdoor attacks. Existing attacks typically rely on fixe…
A Few Words Can Distort Graphs: Knowledge Poisoning Attacks on Graph-based Retrieval-Augmented Generation of Large Language Models
Jiayi Wen, Tianxin Chen, Zhirun Zheng +1
Graph-based Retrieval-Augmented Generation (GraphRAG) has recently emerged as a promising paradigm for enhancing large language models (LLMs) by converting raw text into structured…
PrivDFS: Private Inference via Distributed Feature Sharing against Data Reconstruction Attacks
Zihan Liu, Jiayi Wen, Junru Wu +4
In this paper, we introduce PrivDFS, a distributed feature-sharing framework for input-private inference in image classification. A single holistic intermediate representation in s…