11 papers
SparsePixels: Efficient Convolution for Sparse Data on FPGAs
Ho Fung Tsoi, Dylan Rankin, Vladimir Loncar +1
Inference of standard convolutional neural networks (CNNs) on FPGAs often incurs high latency and a long initiation interval due to the deep nested loops required to densely convol…
da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs
Chang Sun, Zhiqiang Que, Vladimir Loncar +2
Neural networks with a latency requirement on the order of microseconds, like the ones used at the CERN Large Hadron Collider, are typically deployed on FPGAs fully unrolled and pi…
PQuantML: A Tool for End-to-End Hardware-aware Model Compression
Roope Niemi, Anastasiia Petrovych, Arghya Ranjan Das +9
PQuantML is a new open-source, hardware-aware neural network model compression library tailored to end-to-end workflows. Motivated by the need to deploy performant models to enviro…
Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml
Katya Govorkova, Julian Garcia Pardinas, Vladimir Loncar +5
This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which could be used in future high-energy physic…
hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
Jan-Frederik Schulte, Benjamin Ramhorst, Chang Sun +50
We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can b…
wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
Benjamin Hawks, Jason Weitz, Dmitri Demler +13
As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantl…