4 papers
Jack of All Scales: A Versatile FPGA Tensor Block for MXFP Precisions
Marwan Mekhemer, Ahmed Elsousy, Balaji Venkatesh +4
Modern deep learning workloads increasingly rely on narrow numerical formats to improve efficiency and reduce memory footprint. The recently standardized microscaling floating-poin…
A Protocol-Independent Transport Architecture
Kimiya Mohammadtaheri, David Gao, Samuel Zhang +8
The network transport layer is increasingly implemented in the NIC hardware to meet the performance demands of modern workloads, but this has made it difficult to evolve or deploy…
RapidLayout: Fast Hard Block Placement of FPGA-optimized Systolic Arrays using Evolutionary Algorithms
Niansong Zhang, Xiang Chen, Nachiket Kapre
Evolutionary algorithms can outperform conventional placement algorithms such as simulated annealing, analytical placement as well as manual placement on metrics such as runtime, w…
Out-of-Order Dataflow Scheduling for FPGA Overlays
Siddhartha, Nachiket Kapre
We exploit floating-point DSPs in the Arria10 FPGA and multi-pumping feature of the M20K RAMs to build a dataflow-driven soft processor fabric for large graph workloads. In this pa…