2 papers
cs.CR2026
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
Manish Bhatt, Sarthak Munshi, Vineeth Sai Narajala +6
We prove that no continuous, utility-preserving wrapper defense-a function that preprocesses inputs before the model sees them-can make all outputs strictly safe for a…
cs.CR2026
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
Manish Bhatt, Adrian Wood, Idan Habler +1
Production LLM agents with tool-using capabilities require security testing despite their safety training. We adapt Go-Explore to evaluate GPT-4o-mini across 28 experimental runs s…