1 paper
Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu +1
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distri…