1 paper · 1 filter
Chen Ying Claude, Zhihan Luo
Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that survive system prompt replace…