2 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.DC2023★ 2 cited
Miriam: Exploiting Elastic Kernels for Real-time Multi-DNN Inference on Edge GPU
Zhihe Zhao, Neiwen Ling, Nan Guan +1
Many applications such as autonomous driving and augmented reality, require the concurrent running of multiple deep neural networks (DNN) that poses different levels of real-time p…
cs.LG2022★ 2 cited
Moses: Efficient Exploitation of Cross-device Transferable Features for Tensor Program Optimization
Zhihe Zhao, Xian Shuai, Yang Bai +4
Achieving efficient execution of machine learning models has attracted significant attention recently. To generate tensor programs efficiently, a key component of DNN compilers is…