1 paper
Maciej ChrabÄ szcz, Filip Szatkowski, Bartosz Wójcik +3
Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suff…