1 paper
Thiago Sandoval, Ufuk Topcu
Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired pol…