4 papers
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
Kemal Derya, Berk Sunar
Defending large language models (LLMs) against jailbreak attacks, such as Greedy Coordinate Gradient (GCG), remains a challenge, particularly under adaptive threat models where an…
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
Andrew Adiletta, Kathryn Adiletta, Kemal Derya +1
The rapid deployment of Large Language Models (LLMs) has created an urgent need for enhanced security and privacy measures in Machine Learning (ML). LLMs are increasingly being use…
Non-Halting Queries: Exploiting Fixed Points in LLMs
Ghaith Hammouri, Kemal Derya, Berk Sunar
We introduce a new vulnerability that exploits fixed points in autoregressive models and use it to craft queries that never halt. More precisely, for non-halting queries, the LLM n…
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning
M. Caner Tol, Kemal Derya, Berk Sunar
We propose using reinforcement learning to address the challenges of discovering microarchitectural vulnerabilities, such as Spectre and Meltdown, which exploit subtle interactions…