1 paper
Junhyeong Hwangbo, Soohyun Lee, Hyeon Jeon +4
Even LLMs that appear safe during evaluation can still produce harmful responses in deployment. Because stochastic sampling yields different responses to the same prompt, low-proba…