collaborators

7 papers

cs.AI2026

Reasoning Over Space: Enabling Geographic Reasoning for LLM-Based Generative Next POI Recommendation

Dongyi Lv, Qiuyu Ding, Heng-Da Xu +4

Generative recommendation with large language models (LLMs) reframes prediction as sequence generation, yet existing LLM-based recommenders remain limited in leveraging geographic…

cs.DC2026

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection

Yuhang Zhou, Zhibin Wang, Peng Jiang +12

Training large language models faces frequent interruptions due to various faults, demanding robust fault-tolerance. Existing backup-free methods, such as redundant computation, dy…

cs.LG2025

PLATONT: Learning a Platonic Representation for Unified Network Tomography

Chengze Du, Heng Xu, Zhiwei Yu +2

Network tomography aims to infer hidden network states, such as link performance, traffic load, and topology, from external observations. Most existing methods solve these problems…

cs.DC2025

RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training

Heng Xu, Zhiwei Yu, Chengze Du +5

Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to…

cs.DC2025

Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning

Chengze Du, Zhiwei Yu, Heng Xu +3

The rapid growth of large language model (LLM) services imposes increasing demands on distributed GPU inference infrastructure. Most existing scheduling systems follow a reactive p…

cs.NI2025

REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks

Zhiwei Yu, Chengze Du, Heng Xu +3

Community GPU platforms are emerging as a cost-effective and democratized alternative to centralized GPU clusters for AI workloads, aggregating idle consumer GPUs from globally dis…