1 paper
Yubo Wang, Ping Nie, Kai Zou +2
We have witnessed that strong LLMs like Qwen-Math, MiMo, and Phi-4 possess immense reasoning potential inherited from the pre-training stage. With reinforcement learning (RL), thes…