activity
20242026
most citedPyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System

2 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CR2026

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks

Michael Kouremetis, Ads Dawson, Raja Sekhar Rao Dheekonda +1

Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating i…

cs.AI2026

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers

AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current app…

cs.AI2025★ 2 cited

Lessons From Red Teaming 100 Generative AI Products

Blake Bullwinkel, Amanda Minnich, Shiven Chawla +23

In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems. Due to the nascency of the field, there are many open questi…

cs.CR2024★ 2 cited

PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System

Gary D. Lopez Munoz, Amanda J. Minnich, Roman Lutz +17

Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both s…

cs.CL2024★ 2 cited

Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Emman Haider, Daniel Perez-Becker, Thomas Portet +28

Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…