10 papers
Manifold of Failure: Behavioral Attraction Basins in Language Models
Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala +4
While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehensive understanding of AI safety r…
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
Manish Bhatt, Sarthak Munshi, Vineeth Sai Narajala +6
We prove that no continuous, utility-preserving wrapper defense-a function that preprocesses inputs before the model sees them-can make all outputs strictly safe for a…
LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems
Hammad Atta, Ken Huang, Kyriakos Rock Lambros +11
Agentic LLM systems equipped with persistent memory, RAG pipelines, and external tool connectors face a class of attacks - Logic-layer Prompt Control Injection (LPCI) - for which n…
Predictive Coding and Information Bottleneck for Hallucination Detection in Large Language Models
Manish Bhatt
Hallucinations in Large Language Models (LLMs) -- generations that are plausible but factually unfaithful -- remain a critical barrier to high-stakes deployment. Current detection…
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
Manish Bhatt, Adrian Wood, Idan Habler +1
Production LLM agents with tool-using capabilities require security testing despite their safety training. We adapt Go-Explore to evaluate GPT-4o-mini across 28 experimental runs s…
MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm
Vineeth Sai Narajala, Manish Bhatt, Idan Habler +2
The AI trustworthiness crisis threatens to derail the artificial intelligence revolution, with regulatory barriers, security vulnerabilities, and accountability gaps preventing dep…