23 citations · 28 across the 3 of their papers we have counts for
Showing cs.DCShow all
3 papers · 1 filter
cs.DC2023
MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator Systems
Guan Shen, Jieru Zhao, Zeke Wang +5
Along with the fast evolution of deep neural networks, the hardware system is also developing rapidly. As a promising solution achieving high scalability and low manufacturing cost…
cs.DC2023
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs
Hongjing Huang, Yingtao Li, Jie Sun +5
Generalized linear models (GLMs) are a widely utilized family of machine learning models in real-world applications. As data size increases, it is essential to perform efficient di…
cs.DC2020
Optimizing Memory Performance of Xilinx FPGAs under Vitis
Ruoshi Li, Hongjing Huang, Zeke Wang +3
Plenty of research efforts have been devoted to FPGA-based acceleration, due to its low latency and high energy efficiency. However, using the original low-level hardware descripti…