61 citations · 61 across the 4 of their papers we have counts for
5 papers · 1 filter
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
Zhakshylyk Nurlanov, Frank R. Schmidt, Florian Bernard
As Large Language Models (LLMs) are increasingly deployed in safety-critical domains, rigorously evaluating their robustness against adversarial jailbreaks is essential. However, c…
Adaptive Certified Training: Towards Better Accuracy-Robustness Tradeoffs
Zhakshylyk Nurlanov, Frank R. Schmidt, Florian Bernard
As deep learning models continue to advance and are increasingly utilized in real-world systems, the issue of robustness remains a major challenge. Existing certified training meth…
Neural Network Virtual Sensors for Fuel Injection Quantities with Provable Performance Specifications
Eric Wong, Tim Schneider, Joerg Schmitt +2
Recent work has shown that it is possible to learn neural networks with provable guarantees on the output of the model when subject to input perturbations, however these works have…
Wasserstein Adversarial Examples via Projected Sinkhorn Iterations
Eric Wong, Frank R. Schmidt, J. Zico Kolter
A rapidly growing area of work has studied the existence of adversarial examples, datapoints which have been perturbed to fool a classifier, but the vast majority of these works ha…
Scaling provable adversarial defenses
Eric Wong, Frank R. Schmidt, Jan Hendrik Metzen +1
Recent work has developed methods for learning deep network classifiers that are provably robust to norm-bounded adversarial perturbation; however, these methods are currently only…