From the 1 of 12 linked papers with an AI index.
12 papers
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Roy Rinberg, Usha Bhalla, Igor Shilov +2
The paper presents RippleBench, a benchmark that automatically creates multiple‑choice questions about semantically related concepts using a Wikipedia‑based retrieval system, to me…
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Alexander Panfilov, Peter Romov, Igor Shilov +3
We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injecti…
Exploring the limits of strong membership inference attacks on large language models
Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo +13
State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language mode…
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
Igor Shilov, Alex Cloud, Aryo Pradipta Gema +5
Large Language Models increasingly possess capabilities that carry dual-use risks. While data filtering has emerged as a pretraining-time mitigation, it faces significant challenge…
The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models
Euodia Dodd, NataÅ¡a KrÄo, Igor Shilov +1
Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, of…
Counterfactual Influence as a Distributional Quantity
Matthieu Meeus, Igor Shilov, Georgios Kaissis +1
Machine learning models are known to memorize samples from their training data, raising concerns around privacy and generalization. Counterfactual self-influence is a popular metri…