collaborators

6 papers

cs.AI2026

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

Alec Harris, Kasey Corra, Archie Chaudhury +1

Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. H…

cs.AI2026

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Avijit Ghosh, Anka Reuel, Jenny Chim +45

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…

cs.CY2026

Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts

Alexander K. Saeri, Jess Graham, Michael Noetel +185

Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…

cs.AI2025

Shall We Play a Game? Language Models for Open-ended Wargames

Glenn Matlin, Parv Mahajan, Isaac Song +10

LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosin…

cs.LG2025

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3

Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…

cs.AI2025

Patterns and Mechanisms of Contrastive Activation Engineering

Yixiong Hao, Ayush Panda, Stepan Shabalin +1

Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify…