4 papers · 1 filter
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes
Alec Harris, Kasey Corra, Archie Chaudhury +1
Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. H…
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
Shall We Play a Game? Language Models for Open-ended Wargames
Glenn Matlin, Isaac Song, Yixiong Hao +8
LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosin…
Patterns and Mechanisms of Contrastive Activation Engineering
Yixiong Hao, Ayush Panda, Stepan Shabalin +1
Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify…