collaborators

6 papers

cs.DC2026

CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing

Yitao Yuan, Chenqi Zhao, Bohan Zhao +3

Efficiently harnessing GPU compute is critical to improving user experience and reducing operational costs in large language model (LLM) services. However, current inference engine…

cs.DC2025

FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management

Bohan Zhao, Yuanhong Wang, Chenglin Liu +6

Recent developments in large language models (LLMs) have introduced new requirements for efficient and robust training. As LLM clusters scale, node failures, lengthy recoveries, an…

cs.DC2025

SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving

Bohan Zhao, Zane Cao, Yongchao He

As large language models (LLMs) scale out with tensor parallelism (TP) and pipeline parallelism (PP) and production stacks have aggressively optimized the data plane (attention/GEM…

q-fin.ST2025

Kronos: A Foundation Model for the Language of Financial Markets

Yu Shi, Zongliang Fu, Shuo Chen +4

The success of large-scale pre-training paradigm, exemplified by Large Language Models (LLMs), has inspired the development of Time Series Foundation Models (TSFMs). However, their…

cs.DC2025

MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training

Bohan Zhao, Guang Yang, Shuo Chen +4

The rapid escalation in the parameter count of large language models (LLMs) has transformed model training from a single-node endeavor into a highly intricate, cross-node activity.…

cs.DC2025

SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference

Yongchao He, Bohan Zhao, Zheng Cao

As inference workloads for large language models (LLMs) scale to meet growing user demand, pipeline parallelism (PP) has become a widely adopted strategy for multi-GPU deployment,…