collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Prefill Awareness in Large Language Models

Andy Wang, Parv Mahajan, David Demitri Africa +3

Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can reco…

cs.AI2026

Evaluating whether AI models would sabotage AI safety research

Robert Kirk, Alexandra Souly, Kai Fronsdal +2

We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two co…

cs.AI2026

UK AISI Alignment Evaluation Case-Study

Alexandra Souly, Robert Kirk, Jacob Merizian +2

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate…

cs.AI2025

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Andy Zou, Maxwell Lin, Eliot Jones +14

Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these sys…

cs.AI2025

Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction

Michal Bravansky, Vaclav Kubon, Suhas Hariharan +1

Interpreting data is central to modern research. Large language models (LLMs) show promise in providing such natural language interpretations of data, yet simple feature extraction…