Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
How does information access affect LLM monitors' ability to detect sabotage?
Rauno Arike, Raja Mehta Moreno, Rohan Subramani +2
Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control potentially misaligned age…
cs.AI2025
Higher-Order Belief in Incomplete Information MAIDs
Jack Foxabbott, Rohan Subramani, Francis Rhys Ward
Multi-agent influence diagrams (MAIDs) are probabilistic graphical models which represent strategic interactions between agents. MAIDs are equivalent to extensive form games (EFGs)…