activity
20242026
collaborators

9 papers

cs.CR2026

Cordyceps: Covert Control Attacks on LLMs via Data Poisoning

Zedian Shao, Charles Fleming, Teodora Baluta

Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that de…

cs.CV2026

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

Zedian Shao, Hongbin Liu, Yuepeng Hu +1

Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but also raising critical safety and…

cs.CR2025

Watermark Robustness and Radioactivity May Be at Odds in Federated Learning

Leixu Huang, Zedian Shao, Teodora Baluta

Federated learning (FL) enables fine-tuning large language models (LLMs) across distributed data sources. As these sources increasingly include LLM-generated text, provenance track…

cs.CR2025

PromptLocate: Localizing Prompt Injection Attacks

Yuqi Jia, Yupei Liu, Zedian Shao +2

Prompt injection attacks deceive a large language model into completing an attacker-specified task instead of its intended task by contaminating its input data with an injected pro…

cs.LG2025

WebInject: Prompt Injection Attack to Web Agents

Xilong Wang, John Bloch, Zedian Shao +3

Multi-modal large language model (MLLM)-based web agents interact with webpage environments by generating actions based on screenshots of the webpages. In this work, we propose Web…

cs.CR2025

Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment

Zedian Shao, Hongbin Liu, Jaden Mu +1

Prompt injection attack, where an attacker injects a prompt into the original one, aiming to make an Large Language Model (LLM) follow the injected prompt to perform an attacker-ch…