364 citations · 572 across the 21 of their papers we have counts for
1 paper · 2 filters
Yotam Wolf, Noam Wies, Dorin Shteyman +3
Language model alignment has become an important component of AI safety, allowing safe interactions between humans and language models, by enhancing desired behaviors and inhibitin…