1 paper · 1 filter
Anna C. Marbut, Daniel R. Olson, Travis J. Wheeler
Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed requests. Recent work has sho…