3 citations · 3 across the 1 of their papers we have counts for
1 paper
Johannes Treutlein, Dami Choi, Jan Betley +4
One way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data. While this removes the explicit information, implicit i…