4 papers
Persona-Model Collapse in Emergent Misalignment
Davi Bastos Costa, Renato Vicente
Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as emergent misalignment. We pro…
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
Davi Bastos Costa, Felippe Alves, Renato Vicente
Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral resp…
Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia
Davi Bastos Costa, Renato Vicente
Large language models are increasingly deployed in multi-agent settings whose outcomes hinge on social intelligence, motivating evaluations of their interactive capabilities; yet e…
Effective field theory for weakly bound two-neutron halo nuclei: corrections from neutron-neutron effective range
Davi B. Costa, Masaru Hongo, Dam Thanh Son
Using an effective field-theoretical approach, we investigate the properties of weakly bound two-neutron halo nuclei (also known as Borromean nuclei) that do not support a low-ener…