9 papers
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
Sarah Ball, Phil Hackemann
This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious acto…
Automated reproducibility assessments in the social and behavioral sciences using large language models
Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten +7
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can…
Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
Yushi Yang, Shreyansh Padarha, Sarah Ball +2
Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alig…
Don't Walk the Line: Boundary Guidance for Filtered Generation
Sarah Ball, Andreas Haupt
Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probabil…
Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
Sarah Ball, Simeon Allmendinger, Niklas Kühl +2
Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outpu…
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
Sarah Ball, Niki Hasrati, Alexander Robey +4
Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended onto input prompts, elicit disallowed cont…