collaborators

5 papers

cs.DC2026

TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

Lingyun Zhang, Henghua Zhang, Shilei Gu +5

Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training. E…

cs.IR2026

Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders

Zhimin Chen, Chenyu Zhao, Ka Chun Mo +7

Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential mo…

cs.AI2025

Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE

Anxiang Zeng, Haibo Zhang, Hailing Zhang +13

We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…

cs.CL2025

Mid-Training of Large Language Models: A Survey

Kaixiang Mo, Yuxin Shi, Weiwei Weng +4

Large language models (LLMs) are typically developed through large-scale pre-training followed by task-specific fine-tuning. Recent advances highlight the importance of an intermed…

cs.AI2025

Compass-Thinker-7B Technical Report

Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6

Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…