Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Detecting Safety Violations Across Many Agent Traces
Adam Stein, Davis Brown, Hamed Hassani +2
To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex, and sometimes even adversar…
cs.AI2025
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
Sagnik Anupam, Davis Brown, Shuo Li +3
LLM web agents now browse and take actions on the open web, yet current agent evaluations are constrained to sandboxed environments or artificial tasks. We introduce BrowserArena,…