Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Persona-Model Collapse in Emergent Misalignment
Davi Bastos Costa, Renato Vicente
Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as emergent misalignment. We pro…
cs.CL2026
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
Davi Bastos Costa, Felippe Alves, Renato Vicente
Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral resp…