most citedMid-Training of Large Language Models: A Survey

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR

Zijun Min, Bingshuai Liu, Ante Wang +4

Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus…

cs.CL2025

Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings

Pakorn Ueareeworakul, Shuman Liu, Jinghao Feng +7

As global e-commerce rapidly expands into emerging markets, the lack of high-quality semantic representations for low-resource languages has become a decisive bottleneck for retrie…

cs.AI2025

Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE

Anxiang Zeng, Haibo Zhang, Hailing Zhang +13

We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…

cs.AI2025

Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications

Shuyi Xie, Ziqin Liew, Hailing Zhang +5

Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in specialized domains remain underexplored. In e-commerce, existing evaluations-such a…

cs.CL20251 cited

Mid-Training of Large Language Models: A Survey

Kaixiang Mo, Yuxin Shi, Weiwei Weng +4

Large language models (LLMs) are typically developed through large-scale pre-training followed by task-specific fine-tuning. Recent advances highlight the importance of an intermed…

cs.LG2025

SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts

Bingshuai Liu, Ante Wang, Zijun Min +7

Large Language Models (LLMs) increasingly rely on reinforcement learning with verifiable rewards (RLVR) to elicit reliable chain-of-thought reasoning. However, the training process…