1 paper · 1 filter
Suyash Fulay, William Brannon, Shrestha Mohanty +4
Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simul…