3 papers
cs.CR2025
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
Daniel Schwartz, Dmitriy Bespalov, Zhe Wang +2
As large language models (LLMs) become increasingly prevalent, ensuring their robustness against adversarial misuse is crucial. This paper introduces the GAP (Graph of Attacks with…
cs.CR2024
TaeBench: Improving Quality of Toxic Adversarial Examples
Xuan Zhu, Dmitriy Bespalov, Liwen You +2
Toxicity text detectors can be vulnerable to adversarial examples - small perturbations to input text that fool the systems into wrong detection. Existing attack algorithms are tim…
cs.CL2024
Less is More for Improving Automatic Evaluation of Factual Consistency
Tong Wang, Ninad Kulkarni, Yanjun Qi
Assessing the factual consistency of automatically generated texts in relation to source context is crucial for developing reliable natural language generation applications. Recent…