activity
20182026
most citedSleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

39 citations · 110 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Modular Pretraining Enables Access Control

Ethan Roland, Murat Cubuktepe, Erick Martinez +8

AI developers face a dual-use dilemma. An AI capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved with access control, limi…

cs.LG2025

Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs

Igor Shilov, Alex Cloud, Aryo Pradipta Gema +5

Large Language Models increasingly possess capabilities that carry dual-use risks. While data filtering has emerged as a pretraining-time mitigation, it faces significant challenge…

cs.LG20242 cited

Sabotage Evaluations for Frontier Models

Joe Benton, Misha Wagner, Eric Christiansen +13

Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage e…

cs.LG202326 cited

Studying Large Language Model Generalization with Influence Functions

Roger Grosse, Juhan Bae, Cem Anil +14

When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which tr…

cs.LG20223 cited

Path Independent Equilibrium Models Can Better Exploit Test-Time Computation

Cem Anil, Ashwini Pokle, Kaiqu Liang +5

Designing networks capable of attaining better performance with an increased inference budget is important to facilitate generalization to harder problem instances. Recent efforts…

cs.LG2021

Learning to Elect

Cem Anil, Xuchan Bao

Voting systems have a wide range of applications including recommender systems, web search, product design and elections. Limited by the lack of general-purpose analytical tools, i…