1 paper · 1 filter
Youting Wang, Xiao Han, Dingyan Shang +2
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentD…