4 citations · 7 across the 6 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
Valentyn Boreiko, Alexander Panfilov, Vaclav Voracek +2
A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their origi…
cs.LG2022★ 4 cited
Provably Adversarially Robust Nearest Prototype Classifiers
Václav Voráček, Matthias Hein
Nearest prototype classifiers (NPCs) assign to each input point the label of the nearest prototype with respect to a chosen distance metric. A direct advantage of NPCs is that the…
cs.LG2022★ 1 cited
Sound Randomized Smoothing in Floating-Point Arithmetics
Václav Voráček, Matthias Hein
Randomized smoothing is sound when using infinite precision. However, we show that randomized smoothing is no longer sound for limited floating-point precision. We present a simple…