2 papers
cs.CR2026
SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision
Guy Azov, Ofer Rivlin, Guy Shtar
Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or images. Existing defenses often narrow t…
cs.CR2025
ASTRA: Agentic Steerability and Risk Assessment Framework
Itay Hazan, Yael Mathov, Guy Shtar +2
Securing AI agents powered by Large Language Models (LLMs) represents one of the most critical challenges in AI security today. Unlike traditional software, AI agents leverage LLMs…