1 paper · 1 filter
Nikita Kezins, Urbas Ekka, Pascal Berrang +1
Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal guarantees. Providing forma…