activity
20242026
collaborators

5 papers

cs.DC2026

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training

Juntao Zhao, Qi Lu, Wei Jia +13

Modern frameworks for training large foundation models (LFMs) employ dataloaders in a data-parallel manner, with each loader processing a disjoint subset of training data. When pre…

cs.AR2026

Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving

Juntao Zhao, Jiuru Li, Chuan Wu

CPUs are critical for LLM serving due to their availability, cost efficiency, and edge applicability. However, efficient CPU serving is hindered by conflicting prefill/decode resou…

cs.LG2025

QSpec: Speculative Decoding with Complementary Quantization Schemes

Juntao Zhao, Wenhao Lu, Sheng Wang +2

Quantization is widely adopted to accelerate inference and reduce memory consumption in large language models (LLMs). While activation-weight joint quantization enables efficient l…

cs.AI2025

Efficient LLM Serving on Hybrid Real-time and Best-effort Requests

Wan Borui, Zhao Juntao, Jiang Chenyu +2

Recent breakthroughs in large Language Models (LLMs) have enabled various generative tasks on a single model. Real-world services (e.g., OpenAI's ChatGPT [27]) powered by an LLM of…

cs.LG2024

QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices

Juntao Zhao, Borui Wan, Yanghua Peng +3

A number of production deep learning clusters have attempted to explore inference hardware for DNN training, at the off-peak serving hours with many inference GPUs idling. Conducti…