1 paper
Jun Gao, Ziqiang Cao, Shaoyao Huang +2
ChatGPT is instruct-tuned to generate general and human-expected content to align with human preference through Reinforcement Learning from Human Feedback (RLHF), meanwhile resulti…