4 citations · 9 across the 6 of their papers we have counts for
13 papers · 1 filter
Measuring and Mitigating Persona Distortions from AI Writing Assistance
Paul Röttger, Kobi Hackenburg, Hannah Rose Kirk +1
Hundreds of millions of people use artificial intelligence (AI) for writing assistance. Here, we evaluated how AI writing assistance distorts writer personas - their perceived beli…
Diffusion Language Models Are Natively Length-Aware
Vittorio Rossi, Giacomo Cirò, Davide Beltrame +3
Unlike autoregressive language models, which terminate variable-length generation upon predicting an End-of-Sequence (EoS) token, Diffusion Language Models (DLMs) operate over a fi…
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
Flor Miriam Plaza-del-Arco, Paul Röttger, Nino Scherrer +3
Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent wor…
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
Giuseppe Russo, Debora Nozza, Paul Röttger +1
People increasingly rely on Large Language Models (LLMs) for moral advice, which may influence humans' decisions. Yet, little is known about how closely LLMs align with human moral…
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
Dominik Meier, Jan Philip Wahle, Paul Röttger +2
As large language models (LLMs) become integrated into sensitive workflows, concerns grow over their potential to leak confidential information. We propose TrojanStego, a novel thr…
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
Paul Röttger, Musashi Hinck, Valentin Hofmann +4
Large language models (LLMs) are helping millions of users write texts about diverse issues, and in doing so expose users to different ideas and perspectives. This creates concerns…