2 papers
cs.CR2025
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
Daniel Schwartz, Dmitriy Bespalov, Zhe Wang +2
As large language models (LLMs) become increasingly prevalent, ensuring their robustness against adversarial misuse is crucial. This paper introduces the GAP (Graph of Attacks with…
cs.CL2025
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
Aman Goel, Daniel Schwartz, Yanjun Qi
Large language models (LLMs) have demonstrated impressive capabilities across diverse tasks, but they remain susceptible to hallucinations--generating content that appears plausibl…