8 citations · 13 across the 3 of their papers we have counts for
3 papers
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
Andrew Konya, Aviv Ovadya, Kevin Feng +4
We introduce a method to measure the alignment between public will and language model (LM) behavior that can be applied to fine-tuning, online oversight, and pre-release safety che…
Democratic Policy Development using Collective Dialogues and AI
Andrew Konya, Lisa Schirch, Colin Irwin +1
We design and test an efficient democratic process for developing policies that reflect informed public will. The process combines AI-enabled collective dialogues that make deliber…
Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives
Elizabeth Seger, Noemi Dreksler, Richard Moulange +19
Recent decisions by leading AI labs to either open-source their models or to restrict access to their models has sparked debate about whether, and how, increasingly capable AI mode…