activity
20242026
most citedda4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs

3 citations · 3 across the 3 of their papers we have counts for

collaborators
Showing cs.ARShow all

5 papers · 1 filter

cs.AR2026

FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees

Zhiqiang Que, Chang Sun, Haiyang Wang +6

Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or man…

cs.AR2026

HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference

Chang Sun, Zhiqiang Que, Bakhtiar Zadeh +4

Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations directly onto the logic primitive…

cs.AR20263 cited

da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs

Chang Sun, Zhiqiang Que, Vladimir Loncar +2

Neural networks with a latency requirement on the order of microseconds, like the ones used at the CERN Large Hadron Collider, are typically deployed on FPGAs fully unrolled and pi…

cs.AR2025

Enthuse: Efficient Adaptable High-throughput Streaming Aggregation Engines

Philippos Papaphilippou, Wayne Luk

Aggregation queries are a series of computationally-demanding analytics operations on counted, grouped or time series data. They include tasks such as summation or finding the medi…

cs.AR2025

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

Jan-Frederik Schulte, Benjamin Ramhorst, Chang Sun +50

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can b…