5 papers
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
Size Zheng, Xuegui Zheng, Li-wen Chang +1
The exponential growth in Large Language Model (LLM) parameters has transformed model training into an increasingly resource-intensive endeavor. With the stagnation of Moore's Law…
Jano: Adaptive Diffusion Generation with Early-stage Convergence Awareness
Yuyang Chen, Linqian Zeng, Yijin ZHou +2
Diffusion models have achieved remarkable success in generative AI, yet their computational efficiency remains a significant challenge, particularly for Diffusion Transformers (DiT…
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
Chia-chi Hsieh, Zan Zong, Xinyang Chen +3
The growing demand for large language models (LLMs) requires serving systems to handle many concurrent requests with diverse service level objectives (SLOs). This exacerbates head-…
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
Lijuan Jiang, Xingjian Qian, Zhenxiang Ma +4
Pipeline parallelism is an essential distributed parallelism method. Increasingly complex and diverse DNN models necessitate meticulously customized pipeline schedules for performa…
GS-Cache: A GS-Cache Inference Framework for Large-scale Gaussian Splatting Models
Miao Tao, Yuanzhen Zhou, Haoran Xu +10
Rendering large-scale 3D Gaussian Splatting (3DGS) model faces significant challenges in achieving real-time, high-fidelity performance on consumer-grade devices. Fully realizing t…