collaborators

9 papers

cs.AI2026

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

Sarah Ball, Phil Hackemann

This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious acto…

cs.AI2026

Automated reproducibility assessments in the social and behavioral sciences using large language models

Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten +7

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can…

cs.CL2026

Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning

Yushi Yang, Shreyansh Padarha, Sarah Ball +2

Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alig…

cs.LG2026

Don't Walk the Line: Boundary Guidance for Filtered Generation

Sarah Ball, Andreas Haupt

Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probabil…

cs.CY2026

Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting

Sarah Ball, Simeon Allmendinger, Niklas Kühl +2

Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outpu…

cs.CL2025

Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models

Sarah Ball, Niki Hasrati, Alexander Robey +4

Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended onto input prompts, elicit disallowed cont…