26 citations · 76 across the 18 of their papers we have counts for
1 paper · 2 filters
Lev McKinney, Anvith Thudi, Juhan Bae +4
Standard large language model training can create models that produce outputs their trainer deems unacceptable in deployment. The probability of these outputs can be reduced using…