collaborators

5 papers

cs.CR2025

You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors

Bochuan Cao, Changjiang Li, Yuanpu Cao +3

Large language models (LLMs) have been widely adopted across various applications, leveraging customized system prompts for diverse tasks. Facing potential system prompt leakage ri…

cs.CR2025

Your Agent Can Defend Itself against Backdoor Attacks

Li Changjiang, Liang Jiacheng, Cao Bochuan +2

Despite their growing adoption across domains, large language model (LLM)-powered agents face significant security risks from backdoor attacks during training and fine-tuning. Thes…

cs.CL2025

Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation

Yurui Chang, Bochuan Cao, Lu Lin

While large language models have demonstrated exceptional performance across a wide range of tasks, they remain susceptible to hallucinations -- generating plausible yet factually…

cs.CL2025

TruthFlow: Truthful LLM Generation via Representation Flow Correction

Hanyu Wang, Bochuan Cao, Yuanpu Cao +1

Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these m…

cs.CR2024

Data Free Backdoor Attacks

Bochuan Cao, Jinyuan Jia, Chuxuan Hu +5

Backdoor attacks aim to inject a backdoor into a classifier such that it predicts any input with an attacker-chosen backdoor trigger as an attacker-chosen target class. Existing ba…