1 paper · 1 filter
Quan Liu, Han Zhou, Wenquan Wu +2
Ensuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets, where unsafe prompts are paire…