2 papers
cs.CL2025
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
Maluna Menke, Thilo Hagendorff
Large Language Models (LLMs) frequently reproduce the gender- and sexual-identity prejudices embedded in their training corpora, leading to outputs that marginalize LGBTQIA+ users.…
cs.CL2025
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Laurène Vaugrante, Francesca Carlon, Maluna Menke +1
Recent research on large language models (LLMs) has demonstrated their ability to understand and employ deceptive behavior, even without explicit prompting. However, such behavior…