3 citations · 3 across the 3 of their papers we have counts for
5 papers · 1 filter
FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees
Zhiqiang Que, Chang Sun, Haiyang Wang +6
Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or man…
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
Chang Sun, Zhiqiang Que, Bakhtiar Zadeh +4
Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations directly onto the logic primitive…
da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs
Chang Sun, Zhiqiang Que, Vladimir Loncar +2
Neural networks with a latency requirement on the order of microseconds, like the ones used at the CERN Large Hadron Collider, are typically deployed on FPGAs fully unrolled and pi…
Enthuse: Efficient Adaptable High-throughput Streaming Aggregation Engines
Philippos Papaphilippou, Wayne Luk
Aggregation queries are a series of computationally-demanding analytics operations on counted, grouped or time series data. They include tasks such as summation or finding the medi…
hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
Jan-Frederik Schulte, Benjamin Ramhorst, Chang Sun +50
We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can b…