Fan Yin, Philippe Laban, Xiangyu Peng +7
Malicious content generated by large language models (LLMs) can pose varying degrees of harm. Although existing LLM-based moderators can detect harmful content, they struggle to as…