1 paper · 1 filter
Anxiang Zeng, Haibo Zhang, Hailing Zhang +13
We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…