6 citations · 6 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
Tiansheng Huang, Sihao Hu, Ling Liu
The new paradigm of finetuning-as-a-service introduces a new attack surface for Large Language Models (LLMs): a few harmful data uploaded by users can easily trick the finetuning t…
cs.LG2024
Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
Tiansheng Huang, Sihao Hu, Fatih Ilhan +2
Recent studies show that Large Language Models (LLMs) with safety alignment can be jail-broken by fine-tuning on a dataset mixed with harmful data. First time in the literature, we…