1 paper · 1 filter
Mahdi Nazeri, Anne-Kathrin Schmuck, Sadegh Soudjani +1
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new applicati…