2 papers
cs.AI2026
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Agatha Duzan, Asa Cooper Stickland
Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence setti…
cs.SE2025
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
Thomas Kuntz, Agatha Duzan, Hao Zhao +4
Computer use agents are LLM-based agents that can directly interact with a graphical user interface, by processing screenshots or accessibility trees. While these systems are gaini…