5 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Harvey Dam, Jonas Knochelmann, Vinu Joseph +1
We introduce a method to reduce refusal rates of large language models (LLMs) on sensitive content without modifying model weights or prompts. Motivated by the observation that ref…