Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
Lin Li, Qi Zhang, Xander Davies +2
AI is increasingly used to support scientific peer review, from manuscript screening, reviewer assistance to editorial triage. Although such systems promise to reduce reviewer burd…
cs.CL2026
STACK: Adversarial Attacks on LLM Safeguard Pipelines
Ian R. McKenzie, Oskar J. Hollinsworth, Tom Tseng +5
Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic and OpenAI guard their latest Opus 4 model and GPT-5 mode…
cs.CL2024
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
Maximilian Li, Xander Davies, Max Nadeau
Language models often exhibit behaviors that improve performance on a pre-training objective but harm performance on downstream tasks. We propose a novel approach to removing undes…