3 papers
cs.CL2024
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
Jordan Dotzel, Yash Akhauri, Ahmed S. AbouElhamayed +3
Large language models (LLMs) often struggle with strict memory, latency, and power demands. To meet these demands, various forms of dynamic sparsity have been proposed that reduce…
cs.CV2024
Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel
Jordan Dotzel, Bahaa Kotb, James Dotzel +2
Traditional methods, such as JPEG, perform image compression by operating on structural information, such as pixel values or frequency content. These methods are effective to bitra…
cs.AR2023
M4BRAM: Mixed-Precision Matrix-Matrix Multiplication in FPGA Block RAMs
Yuzong Chen, Jordan Dotzel, Mohamed S. Abdelfattah
Mixed-precision quantization is a popular approach for compressing deep neural networks (DNNs). However, it is challenging to scale the performance efficiently with mixed-precision…