2 papers
cs.CR2025
Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
Tianyu Zhang, Zihang Xi, Jingyu Hua +1
In the realm of black-box jailbreak attacks on large language models (LLMs), the feasibility of constructing a narrow safety proxy, a lightweight model designed to predict the atta…
cs.CR2024
The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
Zhili Shen, Zihang Xi, Ying He +3
The rapid adoption of online chatbots represents a significant advancement in artificial intelligence. However, this convenience brings considerable privacy concerns, as prompts ca…