3 papers
cs.LG2025
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
Valentyn Boreiko, Alexander Panfilov, Vaclav Voracek +2
A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their origi…
stat.ML2025
Treatment of Statistical Estimation Problems in Randomized Smoothing for Adversarial Robustness
Vaclav Voracek
Randomized smoothing is a popular certified defense against adversarial attacks. In its essence, we need to solve a problem of statistical estimation which is usually very time-con…
cs.AI2024
Convergence of Some Convex Message Passing Algorithms to a Fixed Point
Vaclav Voracek, Tomas Werner
A popular approach to the MAP inference problem in graphical models is to minimize an upper bound obtained from a dual linear programming or Lagrangian relaxation by (block-)coordi…