Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
LLM-Safety Evaluations Lack Robustness
Tim Beyer, Sophie Xhonneux, Simon Geisler +3
In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of noise, such as small datasets, methodol…
cs.CR2025
Fast Proxies for LLM Robustness Evaluation
Tim Beyer, Jan Schuchardt, Leo Schwinn +1
Evaluating the robustness of LLMs to adversarial attacks is crucial for safe deployment, yet current red-teaming methods are often prohibitively expensive. We compare the ability o…