12 citations · 17 across the 7 of their papers we have counts for
7 papers
Extracting Unlearned Information from LLMs with Activation Steering
Atakan Seyitoğlu, Aleksei Kuvshinov, Leo Schwinn +1
An unintended consequence of the vast pretraining of Large Language Models (LLMs) is the verbatim memorization of fragments of their training data, which may contain sensitive or c…
Caption-Driven Explorations: Aligning Image and Text Embeddings through Human-Inspired Foveated Vision
Dario Zanca, Andrea Zugarini, Simon Dietz +4
Understanding human attention is crucial for vision science and AI. While many models exist for free-viewing, less is known about task-driven image exploration. To address this, we…
Revisiting the Robust Alignment of Circuit Breakers
Leo Schwinn, Simon Geisler
Over the past decade, adversarial training has emerged as one of the few reliable methods for enhancing model robustness against adversarial attacks [Szegedy et al., 2014, Madry et…
Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation
Björn Nieth, Thomas Altstidl, Leo Schwinn +1
Their vulnerability to small, imperceptible attacks limits the adoption of deep learning models to real-world systems. Adversarial training has proven to be one of the most promisi…
Adversarial Attacks and Defenses in Large Language Models: Old and New Threats
Leo Schwinn, David Dobre, Stephan Günnemann +1
Over the past decade, there has been extensive research aimed at enhancing the robustness of neural networks, yet this problem remains vastly unsolved. Here, one major impediment h…
Contrastive Language-Image Pretrained Models are Zero-Shot Human Scanpath Predictors
Dario Zanca, Andrea Zugarini, Simon Dietz +4
Understanding the mechanisms underlying human attention is a fundamental challenge for both vision science and artificial intelligence. While numerous computational models of free-…