41 citations · 164 across the 18 of their papers we have counts for
23 papers · 1 filter
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
Kemal Derya, Berk Sunar
Defending large language models (LLMs) against jailbreak attacks, such as Greedy Coordinate Gradient (GCG), remains a challenge, particularly under adaptive threat models where an…
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
Andrew Adiletta, Kathryn Adiletta, Kemal Derya +1
The rapid deployment of Large Language Models (LLMs) has created an urgent need for enhanced security and privacy measures in Machine Learning (ML). LLMs are increasingly being use…
Rubber Mallet: A Study of High Frequency Localized Bit Flips and Their Impact on Security
Andrew Adiletta, Zane Weissman, Fatemeh Khojasteh Dana +2
The increasing density of modern DRAM has heightened its vulnerability to Rowhammer attacks, which induce bit flips by repeatedly accessing specific memory rows. This paper present…
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
Andrew Adiletta, Berk Sunar
Side-channel attacks on shared hardware resources increasingly threaten confidentiality, especially with the rise of Large Language Models (LLMs). In this work, we introduce Spill…
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning
M. Caner Tol, Kemal Derya, Berk Sunar
We propose using reinforcement learning to address the challenges of discovering microarchitectural vulnerabilities, such as Spectre and Meltdown, which exploit subtle interactions…
FAULT+PROBE: A Generic Rowhammer-based Bit Recovery Attack
Kemal Derya, M. Caner Tol, Berk Sunar
Rowhammer is a security vulnerability that allows unauthorized attackers to induce errors within DRAM cells. To prevent fault injections from escalating to successful attacks, a wi…