12 citations · 14 across the 5 of their papers we have counts for
6 papers
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
Zhewen Yu, Sudarshan Sreeram, Krish Agrawal +6
Deep Neural Networks (DNNs) excel in learning hierarchical representations from raw data, such as images, audio, and text. To compute these DNN models with high performance and ene…
SATAY: A Streaming Architecture Toolflow for Accelerating YOLO Models on FPGA Devices
Alexander Montgomerie-Corcoran, Petros Toupas, Zhewen Yu +1
AI has led to significant advancements in computer vision and image processing tasks, enabling a wide range of applications in real-life scenarios, from autonomous vehicles to medi…
PASS: Exploiting Post-Activation Sparsity in Streaming Architectures for CNN Acceleration
Alexander Montgomerie-Corcoran, Zhewen Yu, Jianyi Cheng +1
With the ever-growing popularity of Artificial Intelligence, there is an increasing demand for more performant and efficient underlying hardware. Convolutional Neural Networks (CNN…
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
Jianyi Cheng, Cheng Zhang, Zhewen Yu +3
Model quantization represents both parameters (weights) and intermediate values (activations) in a more compact format, thereby directly reducing both computational and memory cost…
HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
Petros Toupas, Alexander Montgomerie-Corcoran, Christos-Savvas Bouganis +1
For Human Action Recognition tasks (HAR), 3D Convolutional Neural Networks have proven to be highly effective, achieving state-of-the-art results. This study introduces a novel str…
SAMO: Optimised Mapping of Convolutional Neural Networks to Streaming Architectures
Alexander Montgomerie-Corcoran, Zhewen Yu, Christos-Savvas Bouganis
Significant effort has been placed on the development of toolflows that map Convolutional Neural Network (CNN) models to Field Programmable Gate Arrays (FPGAs) with the aim of auto…