2 papers
cs.LG2025
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
Weilin Cai, Juyong Jiang, Le Qin +3
Expert parallelism has emerged as a key strategy for distributing the computational workload of sparsely-gated mixture-of-experts (MoE) models across multiple devices, enabling the…
cs.IR2025
Improving Sequential Recommendations via Bidirectional Temporal Data Augmentation with Pre-training
Juyong Jiang, Peiyan Zhang, Yingtao Luo +6
Sequential recommendation systems are integral to discerning temporal user preferences. Yet, the task of learning from abbreviated user interaction sequences poses a notable challe…