3 papers
cs.AR2024
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
Xilai Dai, Yuzong Chen, Mohamed S. Abdelfattah
FPGAs offer a flexible platform for accelerating deep neural network (DNN) inference, particularly for non-uniform workloads featuring fine-grained unstructured sparsity and mixed…
cs.AR2023
M4BRAM: Mixed-Precision Matrix-Matrix Multiplication in FPGA Block RAMs
Yuzong Chen, Jordan Dotzel, Mohamed S. Abdelfattah
Mixed-precision quantization is a popular approach for compressing deep neural networks (DNNs). However, it is challenging to scale the performance efficiently with mixed-precision…
cs.AR2023
BRAMAC: Compute-in-BRAM Architectures for Multiply-Accumulate on FPGAs
Yuzong Chen, Mohamed S. Abdelfattah
Deep neural network (DNN) inference using reduced integer precision has been shown to achieve significant improvements in memory utilization and compute throughput with little or n…