1 paper · 1 filter
Dianyun Wang, Qingsen Ma, Yuhu Shang +5
Safety alignment -- training large language models (LLMs) to refuse harmful requests while remaining helpful -- is critical for responsible deployment. Prior work established that…