1 paper · 1 filter
Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis +6
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pr…