1 paper
Yongxi Zhou, Wenbo Ye, Yuanzhe Liu +2
Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported…