57 citations · 83 across the 4 of their papers we have counts for
4 papers
SpotServe: Serving Generative Large Language Models on Preemptible Instances
Xupeng Miao, Chunan Shi, Jiangfei Duan +4
The high computational and memory requirements of generative large language models (LLMs) make it challenging to serve them cheaply. This paper aims to reduce the monetary cost for…
Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
Xupeng Miao, Yujie Wang, Youhe Jiang +4
Transformer models have achieved state-of-the-art performance on various domains of applications and gradually becomes the foundations of the advanced large deep learning (DL) mode…
ZOOMER: Boosting Retrieval on Web-scale Graphs by Regions of Interest
Yuezihan Jiang, Yu Cheng, Hanyu Zhao +6
We introduce ZOOMER, a system deployed at Taobao, the largest e-commerce platform in China, for training and serving GNN-based recommendations over web-scale graphs. ZOOMER is desi…
ROD: Reception-aware Online Distillation for Sparse Graphs
Wentao Zhang, Yuezihan Jiang, Yang Li +6
Graph neural networks (GNNs) have been widely used in many graph-based tasks such as node classification, link prediction, and node clustering. However, GNNs gain their performance…