collaborators

5 papers

cs.CR2026

Evaluating AI Models' Capability to Automate Voice Phishing Attacks

Fred Heiding, Claudio Mayrink Verdun, Simon Lermen +5

Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (…

cs.CR2026

Large-scale online deanonymization with LLMs

Simon Lermen, Daniel Paleka, Joshua Swanson +3

We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer…

cs.CR2025

Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation

Fred Heiding, Simon Lermen

We present an end-to-end demonstration of how attackers can exploit AI safety failures to harm vulnerable populations: from jailbreaking LLMs to generate phishing content, to deplo…

cs.AI2025

Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems

Simon Lermen, Mateusz Dziemian, Natalia Pérez-Campanero Antolín

We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental f…

cs.CR2024

Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects

Fred Heiding, Simon Lermen, Andrew Kao +2

In this paper, we evaluate the capability of large language models to conduct personalized phishing attacks and compare their performance with human experts and AI models from last…