most citedFlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism

2 citations · 3 across the 6 of their papers we have counts for

collaborators
Showing cs.DCShow all

8 papers · 1 filter

cs.DC2025

TridentServe: A Stage-level Serving System for Diffusion Pipelines

Yifei Xia, Fangcheng Fu, Hao Yuan +6

Diffusion pipelines, renowned for their powerful visual generation capabilities, have seen widespread adoption in generative vision tasks (e.g., text-to-image/video). These pipelin…

cs.DC2025

Galvatron: An Automatic Distributed System for Efficient Foundation Model Training

Xinyi Liu, Yujie Wang, Shenhan Zhu +4

Galvatron is a distributed system for efficiently training large-scale Foundation Models. It overcomes the complexities of selecting optimal parallelism strategies by automatically…

cs.DC2025

ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Hao Ge, Junda Feng, Qi Huang +6

Scaling long-context ability is essential for Large Language Models (LLMs). To amortize the memory consumption across multiple devices in long-context training, inter-data partitio…

cs.DC2025

ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments

Youhe Jiang, Fangcheng Fu, Xiaozhe Yao +4

Recent developments in large language models (LLMs) have demonstrated their remarkable proficiency in a range of tasks. Compared to in-house homogeneous GPU clusters, deploying LLM…

cs.DC2025

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs

Youhe Jiang, Fangcheng Fu, Xiaozhe Yao +6

Recent advancements in Large Language Models (LLMs) have led to increasingly diverse requests, accompanied with varying resource (compute and memory) demands to serve them. However…

cs.DC2024

Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment

Haoyang Li, Fangcheng Fu, Sheng Lin +8

To optimize large Transformer model training, both efficient parallel computing and advanced data management are indispensable. However, current methods often assume a stable and u…