Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
Sarah Ball, Phil Hackemann
This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious acto…
cs.AI2026
Automated reproducibility assessments in the social and behavioral sciences using large language models
Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten +7
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can…
cs.AI2025
On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
Sarah Ball, Greg Gluch, Shafi Goldwasser +3
With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with…