Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka +4
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job…
cs.AI2026
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
Gabriele La Malfa, Nitay Alon, Emanuele La Malfa +2
Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly…
cs.AI2026
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
Gabriele La Malfa, Emanuele La Malfa, Saar Cohen +4
Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in a zero-sum game, i.e., where…