1 paper · 1 filter
Jiahao Zhao, Liwei Dong
Unlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as…