6 papers
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes
Alec Harris, Kasey Corra, Archie Chaudhury +1
Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. H…
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…
Shall We Play a Game? Language Models for Open-ended Wargames
Glenn Matlin, Parv Mahajan, Isaac Song +10
LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosin…
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3
Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…
Patterns and Mechanisms of Contrastive Activation Engineering
Yixiong Hao, Ayush Panda, Stepan Shabalin +1
Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify…