2 papers
cs.AI2026
Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
Trilok Padhi, Pinxian Lu, Abdulkadir Erol +5
Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jailbreak research has largely focus…
cs.CR2025
Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems
Gauri Sharma, Vidhi Kulkarni, Miles King +1
Evolving AI systems increasingly deploy multi-agent architectures where autonomous agents collaborate, share information, and delegate tasks through developing protocols. This conn…