1 paper
Jingxin Xu, Guoshun Nan, Sheng Guan +7
Recent AI agents, such as ChatGPT and LLaMA, primarily rely on instruction tuning and reinforcement learning to calibrate the output of large language models (LLMs) with human inte…