9 citations · 17 across the 19 of their papers we have counts for
1 paper · 2 filters
Tiansheng Huang, Sihao Hu, Fatih Ilhan +2
Recent studies show that Large Language Models (LLMs) with safety alignment can be jail-broken by fine-tuning on a dataset mixed with harmful data. First time in the literature, we…