#red teaming

try —

4 papers match

cs.CR2026

GPT-Red: Automated Red Teaming via Self-Play at Scale

Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15

The paper presents GPT-Red, an automated red‑teaming system that uses self‑play to generate novel prompt‑injection attacks against large language models and improve their robustnes…

#red teaming#prompt injection#large language models#self-play
cs.LG2026

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Jiazhen Pan, Bailiang Jian, Paul Hager +19

The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…

#large language models#health AI#safety evaluation#red teaming
cs.CR2026

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

Yi Ting Shen, Kentaroh Toyoda, Alex Leung

The paper introduces AMT‑X, a framework that conducts multi‑turn red‑team attacks on large language models using a phase‑structured state machine and evaluates success with a multi…

#llm safety#red teaming#multi-turn attacks#phase-structured evaluation
cs.CR2026

Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

Xutao Mao, Xiang Zheng, Cong Wang

The paper introduces AHA, an automated system that discovers and documents reusable vulnerability concepts in production LLM agents by hypothesizing, testing, and recording unsafe…

#red teaming#large language model agents#vulnerability discovery#automated testing