7 papers
Reasoning Over Space: Enabling Geographic Reasoning for LLM-Based Generative Next POI Recommendation
Dongyi Lv, Qiuyu Ding, Heng-Da Xu +4
Generative recommendation with large language models (LLMs) reframes prediction as sequence generation, yet existing LLM-based recommenders remain limited in leveraging geographic…
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
Yuhang Zhou, Zhibin Wang, Peng Jiang +12
Training large language models faces frequent interruptions due to various faults, demanding robust fault-tolerance. Existing backup-free methods, such as redundant computation, dy…
PLATONT: Learning a Platonic Representation for Unified Network Tomography
Chengze Du, Heng Xu, Zhiwei Yu +2
Network tomography aims to infer hidden network states, such as link performance, traffic load, and topology, from external observations. Most existing methods solve these problems…
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
Heng Xu, Zhiwei Yu, Chengze Du +5
Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to…
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
Chengze Du, Zhiwei Yu, Heng Xu +3
The rapid growth of large language model (LLM) services imposes increasing demands on distributed GPU inference infrastructure. Most existing scheduling systems follow a reactive p…
REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
Zhiwei Yu, Chengze Du, Heng Xu +3
Community GPU platforms are emerging as a cost-effective and democratized alternative to centralized GPU clusters for AI workloads, aggregating idle consumer GPUs from globally dis…