2 papers
cs.CR2026
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
Ted Kwartler, Alan Aqrawi, Arian Abbasi
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during…
cs.CR2024
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
Ted Kwartler, Nataliia Bagan, Ivan Banny +2
The Single-Turn Crescendo Attack (STCA), first introduced in Aqrawi and Abbasi [2024], is an innovative method designed to bypass the ethical safeguards of text-to-text AI models,…