8 citations · 8 across the 2 of their papers we have counts for
3 papers
GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluat…
LLM Critics Help Catch LLM Bugs
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe +3
Reinforcement learning from human feedback (RLHF) is fundamentally limited by the capacity of humans to correctly evaluate model output. To improve human evaluation ability and ove…
GPT-4 Technical Report
OpenAI, Josh Achiam, Steven Adler +278
We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-wor…