5 papers
Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool
Yan Wang, Xiulong Yuan, Kaiming Yang +16
Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost ope…
Accelerating Compound LLM Training Workloads with Maestro
Xiulong Yuan, Hongqing Chen, Jiaxuan Peng +16
Compound LLM training workloads-such as knowledge distillation and multimodal LLM (MLLM) training-are gaining prominence. These typically comprise heterogeneous components differin…
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
Mengshi Qi, Jiaxuan Peng, Xianlin Zhang +1
3D human pose estimation (3D HPE) has emerged as a prominent research topic, particularly in the realm of RGB-based methods. However, the use of RGB images is often limited by issu…
Synergistic Tensor and Pipeline Parallelism
Mengshi Qi, Jiaxuan Peng, Jie Zhang +3
In the machine learning system, the hybrid model parallelism combining tensor parallelism (TP) and pipeline parallelism (PP) has become the dominant solution for distributed traini…
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
Mengshi Qi, Hao Ye, Jiaxuan Peng +1
Action Quality Assessment (AQA), which aims at automatic and fair evaluation of athletic performance, has gained increasing attention in recent years. However, athletes are often i…