most citedAdversarial Attacks and Defenses in Large Language Models: Old and New Threats

12 citations · 17 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL20241 cited

Extracting Unlearned Information from LLMs with Activation Steering

Atakan Seyitoğlu, Aleksei Kuvshinov, Leo Schwinn +1

An unintended consequence of the vast pretraining of Large Language Models (LLMs) is the verbatim memorization of fragments of their training data, which may contain sensitive or c…

cs.CV2024

Caption-Driven Explorations: Aligning Image and Text Embeddings through Human-Inspired Foveated Vision

Dario Zanca, Andrea Zugarini, Simon Dietz +4

Understanding human attention is crucial for vision science and AI. While many models exist for free-viewing, less is known about task-driven image exploration. To address this, we…

cs.CR2024

Revisiting the Robust Alignment of Circuit Breakers

Leo Schwinn, Simon Geisler

Over the past decade, adversarial training has emerged as one of the few reliable methods for enhancing model robustness against adversarial attacks [Szegedy et al., 2014, Madry et…

cs.LG20241 cited

Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation

Björn Nieth, Thomas Altstidl, Leo Schwinn +1

Their vulnerability to small, imperceptible attacks limits the adoption of deep learning models to real-world systems. Adversarial training has proven to be one of the most promisi…

cs.AI202312 cited

Adversarial Attacks and Defenses in Large Language Models: Old and New Threats

Leo Schwinn, David Dobre, Stephan Günnemann +1

Over the past decade, there has been extensive research aimed at enhancing the robustness of neural networks, yet this problem remains vastly unsolved. Here, one major impediment h…

cs.CV20231 cited

Contrastive Language-Image Pretrained Models are Zero-Shot Human Scanpath Predictors

Dario Zanca, Andrea Zugarini, Simon Dietz +4

Understanding the mechanisms underlying human attention is a fundamental challenge for both vision science and artificial intelligence. While numerous computational models of free-…