8 citations · 8 across the 2 of their papers we have counts for
1 paper · 1 filter
Chenyu Shi, Xiao Wang, Qiming Ge +7
Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer beni…