1 paper
Or Biton, Tomer Krichli, Itai Allouche +1
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. T…