computer security

Adversarial Prompting Framework for AI Safety Assessment

arXiv:2607.13453

summary

The paper introduces an Adversarial Prompting Framework that automatically generates and tests structured adversarial prompts of varying complexity to evaluate the safety and robustness of generative AI models in enterprise settings.

Abstract

Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use of these models may also expose systems to new forms of cyberattacks by different malicious actors -- adversarial prompt attack (APA) being one of the most prominent examples of such threats. This paper presents the implementation of an Adversarial Prompting Framework (APF) for a comprehensive assessment of AI safety. The framework systematically evaluates the resilience of the AI model through the generation of structured adversarial prompts at multiple sophistication levels, from direct harmful requests to advanced encoding-based attacks. Our implementation demonstrates the practical application of this methodology in enterprise environments, providing automated testing capabilities with quantitative security assessment metrics. The results indicate significant variations in the model vulnerabilities across different attack vectors, with encoded prompts presenting the highest success rates in bypassing safety mechanisms.

3 pages, 1 figure, presented as a poster at International Conference on Data Science (CODS), December 17-20, 2025, Pune, India

Topics & keywords

#adversarial prompting#ai safety#generative ai#security assessment#prompt engineering#cybersecurityadversarial prompt attackadversarial prompting frameworkstructured adversarial promptsencoding-based attacksquantitative security metrics
Adversarial Prompting Framework for AI Safety Assessment · wovepaper