1 paper
Andrew Konya, Aviv Ovadya, Kevin Feng +4
We introduce a method to measure the alignment between public will and language model (LM) behavior that can be applied to fine-tuning, online oversight, and pre-release safety che…