collaborators

6 papers

cs.LG2026

Manifold of Failure: Behavioral Attraction Basins in Language Models

Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala +4

While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehensive understanding of AI safety r…

cs.CR2026

The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?

Manish Bhatt, Sarthak Munshi, Vineeth Sai Narajala +6

We prove that no continuous, utility-preserving wrapper defense-a function that preprocesses inputs before the model sees them-can make all outputs strictly safe for a…

cs.CR2026

Adversarial Hubness Detector: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems

Idan Habler, Vineeth Sai Narajala, Stav Koren +2

Retrieval-Augmented Generation (RAG) systems are essential to contemporary AI applications, allowing large language models to obtain external knowledge via vector similarity search…

cs.CR2025

MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm

Vineeth Sai Narajala, Manish Bhatt, Idan Habler +2

The AI trustworthiness crisis threatens to derail the artificial intelligence revolution, with regulatory barriers, security vulnerabilities, and accountability gaps preventing dep…

cs.CR2025

A2AS: Agentic AI Runtime Security and Self-Defense

Eugene Neelou, Ivan Novikov, Max Moroz +15

The A2AS framework is introduced as a security layer for AI agents and LLM-powered applications, similar to how HTTPS secures HTTP. A2AS enforces certified behavior, activates mode…

cs.CR2025

Building A Secure Agentic AI Application Leveraging A2A Protocol

Idan Habler, Ken Huang, Vineeth Sai Narajala +1

As Agentic AI systems evolve from basic workflows to complex multi agent collaboration, robust protocols such as Google's Agent2Agent (A2A) become essential enablers. To foster sec…