1 paper
Iago Alves Brito, Walcy Santos Rezende Rios, Julia Soares Dollis +2
Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under generic categories such as "Identity Hate"…