4 papers
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
Or Zion Eliav, Eyal Lenga, Shir Bernstien +1
Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same tr…
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
Itay Zloczower, Eyal Lenga, Gilad Gressel +1
Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs. Although these models are safety-aligned before release, their safegua…
Who Owns This Agent? Tracing AI Agents Back to Their Owners
Ruben Chocron, Doron Jonathan Ben Chayim, Eyal Lenga +3
AI agents are increasingly deployed to act autonomously in the world, yet there is still no reliable way to trace a harmful agent back to the account that deployed it. This creates…
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower +3
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. Ho…