4 papers
Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs
Vincent Limbach, Jonas Dornbusch, David Lüdke +2
Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense com…
Closing the Distribution Gap in Adversarial Training for LLMs
Chengzhi Hu, Jonas Dornbusch, David Lüdke +2
Adversarial training for LLMs is one of the most promising methods to reliably improve robustness against adversaries. However, despite significant progress, models remain vulnerab…
AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
Tim Beyer, Jonas Dornbusch, Jakob Steimle +3
The rapid expansion of research on Large Language Model (LLM) safety and robustness has produced a fragmented and oftentimes buggy ecosystem of implementations, datasets, and evalu…
A Simple Combination of Diffusion Models for Better Quality Trade-Offs in Image Denoising
Jonas Dornbusch, Emanuel Pfarr, Florin-Alexandru Vasluianu +2
Diffusion models have garnered considerable interest in computer vision, owing both to their capacity to synthesize photorealistic images and to their proven effectiveness in image…