4 papers
NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
Jiajun Hu, Ruthwik Reddy Sunketa, Lei Zhao +3
Recent FPGAs have improved deep learning (DL) inference efficiency through dedicated tensor blocks and in-BRAM computation. ReRAM-based analog in-memory computing (IMC) pushes effi…
ATLAS: Automated HLS for DL-Optimized FPGAs
Ruthwik Reddy Sunketa, Aman Arora
FPGA architectures increasingly incorporate domain-specific in-fabric hardblocks to accelerate DL inference, particularly GEMM, which dominates DL computation. To realize the perfo…
Boosting FPGA Performance with Direct BRAM-DSP Paths
Jiajun Hu, Ruthwik Reddy Sunketa, Andrew Boutros +1
Efficient data movement between memory and compute units is a key performance bottleneck in modern FPGA designs, particularly for deep learning (DL) workloads. In typical FPGA arch…
Programming Domain-Specific FPGA Hardblocks from HLS: An RTL Blackbox Approach
Ruthwik Reddy Sunketa, Jeevesh Choudhury, Aman Arora
Domain-specific Field Programmable Gate Array (FPGA) architectures increasingly integrate specialized hardblocks, such as Tensor Slices, to accelerate artificial intelligence and m…