3 papers
cs.AI2026
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5
Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vul…
cs.RO2026
From Kinematics to Dynamics: Learning to Refine Hybrid Plans for Physically Feasible Execution
Lidor Erez, Shahaf S. Shperberg, Ayal Taitler
In many robotic tasks, agents must traverse a sequence of spatial regions to complete a mission. Such problems are inherently mixed discrete-continuous: a high-level action sequenc…
cs.CR2026
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
Lidor Erez, Omer Hofman, Tamir Nizri +1
Automated LLM vulnerability scanners are increasingly used to assess security risks by measuring different attack type success rates (ASR). Yet the validity of these measurements h…