1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.AI2025
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
Anxiang Zeng, Haibo Zhang, Hailing Zhang +13
We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…
cs.CL2025★ 1 cited
Mid-Training of Large Language Models: A Survey
Kaixiang Mo, Yuxin Shi, Weiwei Weng +4
Large language models (LLMs) are typically developed through large-scale pre-training followed by task-specific fine-tuning. Recent advances highlight the importance of an intermed…
cs.AI2025
Compass-Thinker-7B Technical Report
Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6
Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…