3 papers
cs.AI2026
Prefill Awareness in Large Language Models
Andy Wang, Parv Mahajan, David Demitri Africa +3
Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can reco…
cs.CY2026
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…
cs.CR2026
An Independent Safety Evaluation of Kimi K2.5
Zheng-Xin Yong, Parv Mahajan, Andy Wang +12
Kimi K2.5 is an open-weight LLM that rivals closed models across coding, multimodal, and agentic benchmarks, but was released without an accompanying safety evaluation. In this wor…