1 paper · 1 filter
Vincent Conitzer, Rachel Freedman, Jobst Heitzig +9
Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tu…