Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
Filippo Tonini, Federico Torrielli, Anton Danholt Lautrup +3
As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisions together: discussing, negotiating, and acting on shared tas…
cs.AI2025
SHARPIE: A Modular Framework for Reinforcement Learning and Human-AI Interaction Experiments
Hüseyin Aydın, Kevin Godin-Dubois, Libio Goncalvez Braz +6
Reinforcement learning (RL) offers a general approach for modeling and training AI agents, including human-AI interaction scenarios. In this paper, we propose SHARPIE (Shared Human…