13 citations · 24 across the 2 of their papers we have counts for
2 papers
cs.AI2024★ 11 cited
A Safe Harbor for AI Evaluation and Red Teaming
Shayne Longpre, Sayash Kapoor, Kevin Klyman +20
Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by promi…
cs.LG2023★ 13 cited
Black Box Adversarial Prompting for Foundation Models
Natalie Maus, Patrick Chao, Eric Wong +1
Prompting interfaces allow users to quickly adjust the output of generative models in both vision and language. However, small changes and design choices in the prompt can lead to…