5.3k citations · 12.9k across the 24 of their papers we have counts for
Showing cs.CYShow all
2 papers · 1 filter
cs.CY2025★ 4 cited
OpenAI's Approach to External Red Teaming for AI Models and Systems
Lama Ahmad, Sandhini Agarwal, Michael Lampe +1
Red teaming has emerged as a critical practice in assessing the possible risks of AI models and systems. It aids in the discovery of novel risks, stress testing possible gaps in ex…
cs.CY2024★ 3 cited
First-Person Fairness in Chatbots
Tyna Eloundou, Alex Beutel, David G. Robinson +7
Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-maki…