23 citations · 23 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 23 cited
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mantas Mazeika, Long Phan, Xuwang Yin +9
Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a sta…
cs.CL2023
Learning Globally Optimized Language Structure via Adversarial Training
Xuwang Yin
Recent work has explored integrating autoregressive language models with energy-based models (EBMs) to enhance text generation capabilities. However, learning effective EBMs for te…