13 citations · 13 across the 1 of their papers we have counts for
3 papers
cs.DC2018
Interstellar: Using Halide's Scheduling Language to Analyze DNN Accelerators
Xuan Yang, Mingyu Gao, Qiaoyi Liu +9
We show that DNN accelerator micro-architectures and their program mappings represent specific choices of loop order and hardware parallelism for computing the seven nested loops o…
cs.AR2016★ 13 cited
FPMax: a 106GFLOPS/W at 217GFLOPS/mm2 Single-Precision FPU, and a 43.7GFLOPS/W at 74.6GFLOPS/mm2 Double-Precision FPU, in 28nm UTBB FDSOI
Jing Pu, Sameh Galal, Xuan Yang +2
FPMax implements four FPUs optimized for latency or throughput workloads in two precisions, fabricated in 28nm UTBB FDSOI. Each unit's parameters, e.g pipeline stages, booth encodi…
cs.DC2016
A Systematic Approach to Blocking Convolutional Neural Networks
Xuan Yang, Jing Pu, Blaine Burton Rister +6
Convolutional Neural Networks (CNNs) are the state of the art solution for many computer vision problems, and many researchers have explored optimized implementations. Most impleme…