1 paper
Zhihao Ding, Jinming Li, Ze Lu +1
Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed binary classification task, impl…