activity
20182026
most citedThe Serverless Computing Survey: A Technical Primer for Design Architecture

179 citations · 198 across the 12 of their papers we have counts for

collaborators
Showing cs.DCShow all

8 papers · 1 filter

cs.DC2026

SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity

Zhenghao Gan, Yichen Bao, Yifei Liu +3

Efficient LLM inference scheduling is crucial for user experience. However, LLM inferences exhibit remarkable demand uncertainty (with unknown output length beforehand) and hybridi…

cs.DC20223 cited

VELTAIR: Towards High-Performance Multi-tenant Deep Learning Services via Adaptive Compilation and Scheduling

Zihan Liu, Jingwen Leng, Zhihui Zhang +3

Deep learning (DL) models have achieved great success in many application domains. As such, many industrial companies such as Google and Facebook have acknowledged the importance o…

cs.DC2022179 cited

The Serverless Computing Survey: A Technical Primer for Design Architecture

Zijun Li, Linsong Guo, Jiagan Cheng +3

The development of cloud infrastructures inspires the emergence of cloud-native computing. As the most promising architecture for deploying microservices, serverless computing has…

cs.DC20215 cited

Characterizing and Demystifying the Implicit Convolution Algorithm on Commercial Matrix-Multiplication Accelerators

Yangjie Zhou, Mengtian Yang, Cong Guo +5

Many of today's deep neural network accelerators, e.g., Google's TPU and NVIDIA's tensor core, are built around accelerating the general matrix multiplication (i.e., GEMM). However…

cs.DC20201 cited

DLFusion: An Auto-Tuning Compiler for Layer Fusion on Deep Neural Network Accelerator

Zihan Liu, Jingwen Leng, Quan Chen +4

Many hardware vendors have introduced specialized deep neural networks (DNN) accelerators owing to their superior performance and efficiency. As such, how to generate and optimize…

cs.DC20201 cited

Towards QoS-Aware and Resource-Efficient GPU Microservices Based on Spatial Multitasking GPUs In Datacenters

Wei Zhang, Quan Chen, Kaihua Fu +6

While prior researches focus on CPU-based microservices, they are not applicable for GPU-based microservices due to the different contention patterns. It is challenging to optimize…