1 paper · 1 filter
Weitao Feng, Lixu Wang, Peizhuo Lv +5
As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studies assume that attackers rely on superv…