1 paper
Yunze Xiao, Tingyu He, Lionel Z. Wang +6
This paper introduces JiraiBench, the first bilingual benchmark for evaluating large language models' effectiveness in detecting self-destructive content across Chinese and Japanes…