activity
20242026
collaborators
Showing 2024Show all

9 papers · 1 filter

cs.CL2024

Best-of-N Jailbreaking

John Hughes, Sara Price, Aengus Lynch +7

We introduce Best-of-N (BoN) Jailbreaking, a simple black-box algorithm that jailbreaks frontier AI systems across modalities. BoN Jailbreaking works by repeatedly sampling variati…

cs.CL2024

Failures to Find Transferable Image Jailbreaks Between Vision-Language Models

Rylan Schaeffer, Dan Valentine, Luke Bailey +12

The integration of new modalities into frontier AI systems offers exciting capabilities, but also increases the possibility such systems can be adversarially manipulated in undesir…

cs.LG2024

Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach

Tony T. Wang, John Hughes, Henry Sleight +7

Defending large language models against jailbreaks so that they never engage in a broadly-defined set of forbidden behaviors is an open problem. In this paper, we investigate the d…

cs.CV2024

Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Frontier AI Models

Sunny Duan, Mikail Khona, Abhiram Iyer +2

Frontier AI systems are making transformative impacts across society, but such benefits are not without costs: models trained on web-scale datasets containing personal and private…

cs.LG2024

In-Context Learning of Energy Functions

Rylan Schaeffer, Mikail Khona, Sanmi Koyejo

In-context learning is a powerful capability of certain machine learning models that arguably underpins the success of today's frontier AI models. However, in-context learning is c…

cs.LG2024

Quantifying Variance in Evaluation Benchmarks

Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer +5

Evaluation benchmarks are the cornerstone of measuring capabilities of large language models (LLMs), as well as driving progress in said capabilities. Originally designed to make c…