3 papers
cs.LG2026
Mechanistically Eliciting Latent Behaviors in Language Models
Andrew Mack, Nina Panickssery, Alexander Matt Turner
We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systemat…
cs.LG2025
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
Igor Shilov, Alex Cloud, Aryo Pradipta Gema +5
Large Language Models increasingly possess capabilities that carry dual-use risks. While data filtering has emerged as a pretraining-time mitigation, it faces significant challenge…
cs.LG2025
Mitigating Many-Shot Jailbreaking
Christopher M. Ackerman, Nina Panickssery
Many-shot jailbreaking (MSJ) is an adversarial technique that exploits the long context windows of modern LLMs to circumvent model safety training by including in the prompt many e…