1 paper
Mahesh Kumar Nandwana, Youngwan Lim, Joseph Liu +3
Large Language Models (LLMs) are typically aligned for safety during the post-training phase; however, they may still generate inappropriate outputs that could potentially pose ris…