3 citations · 5 across the 18 of their papers we have counts for
1 paper · 2 filters
Neel Jain, Aditya Shrivastava, Chenyang Zhu +6
A key component of building safe and reliable language models is enabling the models to appropriately refuse to follow certain instructions or answer certain questions. We may want…