3 papers
cs.AI2026
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka +4
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job…
cs.AI2026
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
Gabriele La Malfa, Emanuele La Malfa, Saar Cohen +4
Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in a zero-sum game, i.e., where…
cs.MA2025
Fairness Aware Reinforcement Learning via Proximal Policy Optimization
Gabriele La Malfa, Jie M. Zhang, Michael Luck +1
Fairness in multi-agent systems (MAS) focuses on equitable reward distribution among agents in scenarios involving sensitive attributes such as race, gender, or socioeconomic statu…