1 paper
Aryan Dutt, Rui Mao, Anupam Chattopadhyay
We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in m…