6 papers
LLM-Safety Evaluations Lack Robustness
Tim Beyer, Sophie Xhonneux, Simon Geisler +3
In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of noise, such as small datasets, methodol…
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
Leo Schwinn, Moritz Ladenburger, Tim Beyer +3
Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, thes…
Sampling-aware Adversarial Attacks Against Large Language Models
Tim Beyer, Yan Scholten, Leo Schwinn +1
To guarantee safe and robust deployment of large language models (LLMs) at scale, it is critical to accurately assess their adversarial robustness. Existing adversarial attacks typ…
AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
Tim Beyer, Jonas Dornbusch, Jakob Steimle +3
The rapid expansion of research on Large Language Model (LLM) safety and robustness has produced a fragmented and oftentimes buggy ecosystem of implementations, datasets, and evalu…
Jailbreak Distillation: Renewable Safety Benchmarking
Jingyu Zhang, Ahmed Elgohary, Xiawei Wang +5
Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a no…
Fast Proxies for LLM Robustness Evaluation
Tim Beyer, Jan Schuchardt, Leo Schwinn +1
Evaluating the robustness of LLMs to adversarial attacks is crucial for safe deployment, yet current red-teaming methods are often prohibitively expensive. We compare the ability o…