3 papers
cs.LG2026
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
Bingshuai Liu, Ante Wang, Zijun Min +7
Large Language Models (LLMs) increasingly rely on reinforcement learning with verifiable rewards (RLVR) to elicit reliable chain-of-thought reasoning. However, the training process…
cs.CL2025
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
Pakorn Ueareeworakul, Shuman Liu, Jinghao Feng +7
As global e-commerce rapidly expands into emerging markets, the lack of high-quality semantic representations for low-resource languages has become a decisive bottleneck for retrie…
cs.AI2025
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
Anxiang Zeng, Haibo Zhang, Hailing Zhang +13
We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…