10 papers
FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees
Zhiqiang Que, Chang Sun, Haiyang Wang +6
Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or man…
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
Chang Sun, Zhiqiang Que, Bakhtiar Zadeh +4
Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations directly onto the logic primitive…
da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs
Chang Sun, Zhiqiang Que, Vladimir Loncar +2
Neural networks with a latency requirement on the order of microseconds, like the ones used at the CERN Large Hadron Collider, are typically deployed on FPGAs fully unrolled and pi…
Low-Latency FPGA Control System for Real-Time Neural Network Processing in CCD-Based Trapped-Ion Qubit Measurement
Binglei Lou, Gautham Duddi Krishnaswaroop, Filip Wojcicki +5
Accurate and low-latency qubit state measurement is critical for trapped-ion quantum computing. While deep neural networks (DNNs) have been integrated to enhance detection fidelity…
JetFormer: A Scalable and Efficient Transformer for Jet Tagging from Offline Analysis to FPGA Triggers
Ruoqing Zheng, Chang Sun, Qibin Liu +7
We present JetFormer, a versatile and scalable encoder-only Transformer architecture for particle jet tagging at the Large Hadron Collider (LHC). Unlike prior approaches that are o…
Enthuse: Efficient Adaptable High-throughput Streaming Aggregation Engines
Philippos Papaphilippou, Wayne Luk
Aggregation queries are a series of computationally-demanding analytics operations on counted, grouped or time series data. They include tasks such as summation or finding the medi…