Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
Lidor Erez, Omer Hofman, Tamir Nizri +1
Automated LLM vulnerability scanners are increasingly used to assess security risks by measuring different attack type success rates (ASR). Yet the validity of these measurements h…
cs.CR2026
AgenTRIM: Tool Risk Mitigation for Agentic AI
Roy Betser, Amit Giloni, Shamik Bose +4
AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks. While such tools extend capability, improper tool permissions introduce security risk…