Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Model Spec Midtraining: Improving How Alignment Training Generalizes
Chloe Li, Nevan Wichers, Sara Price +2
Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- trai…
cs.AI2026
Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
Chloe Li, Mary Phuong, Daniel Tan
As AI systems become more capable of complex agentic tasks, they also become more capable of pursuing undesirable objectives and causing harm. Previous work has attempted to catch…