3 papers
cs.CR2026
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
Ted Kwartler, Alan Aqrawi, Arian Abbasi
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during…
cs.CR2024
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
Ted Kwartler, Nataliia Bagan, Ivan Banny +2
The Single-Turn Crescendo Attack (STCA), first introduced in Aqrawi and Abbasi [2024], is an innovative method designed to bypass the ethical safeguards of text-to-text AI models,…
cs.CR2024
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation
Ted Kwartler, Matthew Berman, Alan Aqrawi
This study explores the ability of Large Language Model (LLM) agents to detect and correct hallucinations in AI-generated content. A primary agent was tasked with creating a blog a…