1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Dalia Ali, Dora Zhao, Allison Koenecke +1
Although large language models (LLMs) are increasingly trained using human feedback for safety and alignment with human values, alignment decisions often overlook human social dive…