2 papers
cs.LG2026
Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values
Henry Bell, Lara Neubauer da Costa Schertel, Bochu Ding +1
A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignme…
cs.CL2026
Reflect: Transparent Principle-Guided Reasoning for Constitutional Alignment at Scale
Henry Bell, Caroline Zhang, Mohammed Mobasserul Haque +3
The constitutional framework of alignment aims to align large language models (LLMs) with value-laden principles written in natural language (such as to avoid using biased language…