activity
20212024
most citedAllo: A Programming Model for Composable Accelerator Design

40 citations · 51 across the 9 of their papers we have counts for

collaborators

9 papers

cs.LG20242 cited

Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on Circuits

Chenhui Deng, Zichao Yue, Cunxi Yu +4

While graph neural networks (GNNs) have gained popularity for learning circuit representations in various electronic design automation (EDA) tasks, they face challenges in scalabil…

cs.CL2024

Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models

Jordan Dotzel, Yash Akhauri, Ahmed S. AbouElhamayed +3

Large language models (LLMs) often struggle with strict memory, latency, and power demands. To meet these demands, various forms of dynamic sparsity have been proposed that reduce…

cs.PL202440 cited

Allo: A Programming Model for Composable Accelerator Design

Hongzheng Chen, Niansong Zhang, Shaojie Xiang +3

Special-purpose hardware accelerators are increasingly pivotal for sustaining performance improvements in emerging applications, especially as the benefits of technology scaling co…

cs.CL20247 cited

UniSparse: An Intermediate Language for General Sparse Format Customization

Jie Liu, Zhongyuan Zhao, Zijian Ding +3

The ongoing trend of hardware specialization has led to a growing use of custom data formats when processing sparse workloads, which are typically memory-bound. These formats facil…

cs.CV2024

Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel

Jordan Dotzel, Bahaa Kotb, James Dotzel +2

Traditional methods, such as JPEG, perform image compression by operating on structural information, such as pixel values or frequency content. These methods are effective to bitra…

cs.LG20242 cited

Trainable Fixed-Point Quantization for Deep Learning Acceleration on FPGAs

Dingyi Dai, Yichi Zhang, Jiahao Zhang +4

Quantization is a crucial technique for deploying deep learning models on resource-constrained devices, such as embedded FPGAs. Prior efforts mostly focus on quantizing matrix mult…