#red teaming
4 papers match
GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15
The paper presents GPT-Red, an automated red‑teaming system that uses self‑play to generate novel prompt‑injection attacks against large language models and improve their robustnes…
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Jiazhen Pan, Bailiang Jian, Paul Hager +19
The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Yi Ting Shen, Kentaroh Toyoda, Alex Leung
The paper introduces AMT‑X, a framework that conducts multi‑turn red‑team attacks on large language models using a phase‑structured state machine and evaluates success with a multi…
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Xutao Mao, Xiang Zheng, Cong Wang
The paper introduces AHA, an automated system that discovers and documents reusable vulnerability concepts in production LLM agents by hypothesizing, testing, and recording unsafe…