179 citations · 197 across the 9 of their papers we have counts for
11 papers
VELTAIR: Towards High-Performance Multi-tenant Deep Learning Services via Adaptive Compilation and Scheduling
Zihan Liu, Jingwen Leng, Zhihui Zhang +3
Deep learning (DL) models have achieved great success in many application domains. As such, many industrial companies such as Google and Facebook have acknowledged the importance o…
The Serverless Computing Survey: A Technical Primer for Design Architecture
Zijun Li, Linsong Guo, Jiagan Cheng +3
The development of cloud infrastructures inspires the emergence of cloud-native computing. As the most promising architecture for deploying microservices, serverless computing has…
Characterizing and Demystifying the Implicit Convolution Algorithm on Commercial Matrix-Multiplication Accelerators
Yangjie Zhou, Mengtian Yang, Cong Guo +5
Many of today's deep neural network accelerators, e.g., Google's TPU and NVIDIA's tensor core, are built around accelerating the general matrix multiplication (i.e., GEMM). However…
Dubhe: Towards Data Unbiasedness with Homomorphic Encryption in Federated Learning Client Selection
Shulai Zhang, Zirui Li, Quan Chen +3
Federated learning (FL) is a distributed machine learning paradigm that allows clients to collaboratively train a model over their own local data. FL promises the privacy of client…
DLFusion: An Auto-Tuning Compiler for Layer Fusion on Deep Neural Network Accelerator
Zihan Liu, Jingwen Leng, Quan Chen +4
Many hardware vendors have introduced specialized deep neural networks (DNN) accelerators owing to their superior performance and efficiency. As such, how to generate and optimize…
How Far Does BERT Look At:Distance-based Clustering and Analysis of BERTs Attention
Yue Guan, Jingwen Leng, Chao Li +2
Recent research on the multi-head attention mechanism, especially that in pre-trained models such as BERT, has shown us heuristics and clues in analyzing various aspects of the mec…