3 papers
cs.DC2026
TerraceMoE: A Cost Model for Hierarchical MoE All-to-All Communication
Weicheng Xue, Bingqiang Wang, Li Yuan +2
Hierarchical two-hop dispatch can reduce slow-fabric traffic in expert-parallel Mixture-of-Experts training, but it adds a second collective and an arrival-side operator chain. We…
cs.CL2025
GPT as a Monte Carlo Language Tree: A Probabilistic Perspective
Kun-Peng Ning, Jia-Yu Yao, Yu-Yang Liu +2
Large Language Models (LLMs), such as GPT, are considered to learn the latent distributions within large-scale web-crawl datasets and accomplish natural language processing (NLP) t…
cs.LG2024
Is Parameter Collision Hindering Continual Learning in LLMs?
Shuo Yang, Kun-Peng Ning, Yu-Yang Liu +4
Large Language Models (LLMs) often suffer from catastrophic forgetting when learning multiple tasks sequentially, making continual learning (CL) essential for their dynamic deploym…