3 papers
cs.AI2026
OSGuard: A Benchmark for Safety in Computer-Use Agents
Mina Mohammadmirzaei, Jeffrey Flanigan
Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks. However, task success alone can miss failures in which an agent reaches the…
cs.CL2026
Right or Wrong, Models Comply: Directional Blindness in LLM Moral Judgment
Jihye Kim, Jeffrey Flanigan
As language models take integrated roles across many domains, the response of LLMs to user pushback becomes a critical alignment property. Yet many existing evaluations treat compl…
cs.CL2026
Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents
Brendan King, Jeffrey Flanigan
AI coding agents have rapidly transformed software engineering, powering widely used interactive coding assistants. Despite their interactive real-world use, existing benchmarks ev…