35 citations · 36 across the 12 of their papers we have counts for
1 paper · 1 filter
Danting Zhang, Bei Peng, Robert Loftin
Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a…