1 paper
Dylan Feng, Pragya Srivastava, Anca Dragan +1
Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response patterns that are unforeseen by mod…