4 papers
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5
Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vul…
MAPS: A Multilingual Benchmark for Agent Performance and Security
Omer Hofman, Jonathan Brokman, Oren Rachmil +7
Agentic AI systems, which build on Large Language Models (LLMs) and interact with tools and memory, have rapidly advanced in capability and scope. Yet, since LLMs have been shown t…
Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
Oren Rachmil, Avishag Shapira, Roy Betser +5
As organizations increasingly deploy LLMs in sensitive domains such as legal, financial, and medical settings, ensuring alignment with internal organizational policies has become a…
Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis
Jonathan Brokman, Omer Hofman, Oren Rachmil +6
This report presents a comparative analysis of open-source vulnerability scanners for conversational large language models (LLMs). As LLMs become integral to various applications,…