1 paper · 1 filter
Wenliang Shan, Michael Fu, Rui Yang +1
Safety alignment is critical for LLM-powered systems. While recent LLM-powered guardrail approaches such as LlamaGuard achieve high detection accuracy of unsafe inputs written in E…