2 papers
cs.LG2026
Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance
Jiawen Zhang, Lipeng He, Kejia Chen +4
Fine-tuning safety-aligned large language models (LLMs) can substantially compromise their safety. Previous approaches require many safety samples or calibration sets, which not on…
cs.CR2025
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
Jiawen Zhang, Kejia Chen, Zunlei Feng +4
With the growing popularity of LLMs among the general public users, privacy-preserving and adversarial robustness have become two pressing demands for LLM-based services, which hav…