1 paper · 1 filter
Neel Jain, Aditya Shrivastava, Chenyang Zhu +6
A key component of building safe and reliable language models is enabling the models to appropriately refuse to follow certain instructions or answer certain questions. We may want…