1 paper
Luoming Hu, Jingjie Zeng, Liang Yang +1
Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the in…