5 papers
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka +4
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job…
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
Gabriele La Malfa, Emanuele La Malfa, Saar Cohen +4
Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in a zero-sum game, i.e., where…
Large Language Models Miss the Multi-Agent Mark
Emanuele La Malfa, Gabriele La Malfa, Samuele Marro +5
Recent interest in Multi-Agent Systems of Large Language Models (MAS LLMs) has led to an increase in frameworks leveraging multiple LLMs to tackle complex tasks. However, much of t…
Fairness Aware Reinforcement Learning via Proximal Policy Optimization
Gabriele La Malfa, Jie M. Zhang, Michael Luck +1
Fairness in multi-agent systems (MAS) focuses on equitable reward distribution among agents in scenarios involving sensitive attributes such as race, gender, or socioeconomic statu…
Using Protected Attributes to Consider Fairness in Multi-Agent Systems
Gabriele La Malfa, Jie M. Zhang, Michael Luck +1
Fairness in Multi-Agent Systems (MAS) has been extensively studied, particularly in reward distribution among agents in scenarios such as goods allocation, resource division, lotte…