1 paper · 1 filter
Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11
Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follows harmful instructions. Whe…