Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
Iago Alves Brito, Walcy Santos Rezende Rios, Julia Soares Dollis +2
Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under generic categories such as "Identity Hate"…
cs.CL2026
ToxSyn-PT: A Synthetic Fine-Grained Dataset of Minority-Targeted Toxic Language in Portuguese
Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Farber +2
The development of robust hate speech detection systems remains limited by the lack of large-scale, fine-grained training data, especially for languages beyond English. Existing co…