23 citations · 40 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 23 cited
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mantas Mazeika, Long Phan, Xuwang Yin +9
Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a sta…
cs.CL2024★ 4 cited
PAL: Proxy-Guided Black-Box Attack on Large Language Models
Chawin Sitawarin, Norman Mu, David Wagner +1
Large Language Models (LLMs) have surged in popularity in recent months, but they have demonstrated concerning capabilities to generate harmful content when manipulated. While tech…
cs.CV2021★ 13 cited
SLIP: Self-supervision meets Language-Image Pre-training
Norman Mu, Alexander Kirillov, David Wagner +1
Recent work has shown that self-supervised pre-training leads to improvements over supervised learning on challenging visual recognition tasks. CLIP, an exciting new approach to le…