3 papers
cs.DC2026
TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training
Lingyun Zhang, Henghua Zhang, Shilei Gu +5
Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training. E…
cs.AI2025
LoongFlow: Directed Evolutionary Search via a Cognitive Plan-Execute-Summarize Paradigm
Chunhui Wan, Xunan Dai, Zhuo Wang +5
The transition from static Large Language Models (LLMs) to self-improving agents is hindered by the lack of structured reasoning in traditional evolutionary approaches. Existing me…
cs.DC2025
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
Jian Tian, Shuailong Li, Yang Cao +8
The evolution of Large Language Model (LLM) serving towards complex, distributed architectures--specifically the P/D-separated, large-scale DP+EP paradigm--introduces distinct sche…