From the 1 of 4 linked papers with an AI index.
1 paper · 1 filter
Ohjoon Kwon, Donghyeon Jeon, Nayoung Choi +6
Most prior safety research of large language models (LLMs) has focused on enhancing the alignment of LLMs to better suit the safety requirements of humans. However, internalizing s…