activity
20242026
most citedLessons From Red Teaming 100 Generative AI Products

2 citations · 6 across the 10 of their papers we have counts for

collaborators
Showing cs.CRShow all

5 papers · 1 filter

cs.CR2026

Evading Chain-of-Thought Monitoring Through Model Poisoning

Giorgio Severi, Shujaat Mirza, Blake Bullwinkel +1

Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its ac…

cs.CR2026

The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers

Blake Bullwinkel, Giorgio Severi, Keegan Hines +3

Detecting whether a model has been poisoned is a longstanding problem in AI security. In this work, we present a practical scanner for identifying sleeper agent-style backdoors in…

cs.CR2025

A Systematization of Security Vulnerabilities in Computer Use Agents

Daniel Jones, Giorgio Severi, Martin Pouliot +7

Computer Use Agents (CUAs), autonomous systems that interact with software interfaces via browsers or virtual machines, are rapidly being deployed in consumer and enterprise enviro…

cs.CR2025

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks

Blake Bullwinkel, Mark Russinovich, Ahmed Salem +8

Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and…

cs.CR2024★ 2 cited

PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System

Gary D. Lopez Munoz, Amanda J. Minnich, Roman Lutz +17

Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both s…