1 paper
Chen Zheng, Yiyuan Ma, Yuan Yang +11
The development of alignment and reasoning capabilities in large language models has seen remarkable progress through two paradigms: instruction tuning and reinforcement learning f…