4 papers
Kolmogorov--Arnold Networks for Small Language Models
Felippe Alves, Renato Vicente
Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alte…
Persona-Model Collapse in Emergent Misalignment
Davi Bastos Costa, Renato Vicente
Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as emergent misalignment. We pro…
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
Davi Bastos Costa, Felippe Alves, Renato Vicente
Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral resp…
Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia
Davi Bastos Costa, Renato Vicente
Large language models are increasingly deployed in multi-agent settings whose outcomes hinge on social intelligence, motivating evaluations of their interactive capabilities; yet e…