Showing cs.ARShow all
3 papers · 1 filter
cs.AR2025
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
Jude Haris, José Cano
Large Language Models (LLMs) have become increasingly prominent for daily tasks, from improving sound-totext translation to generating additional frames for the latest video games.…
cs.AR2025
Accelerating Transposed Convolutions on FPGA-based Edge Devices
Jude Haris, José Cano
Transposed Convolutions (TCONV) enable the up-scaling mechanism within generative Artificial Intelligence (AI) models. However, the predominant Input-Oriented Mapping (IOM) method…
cs.AR2024
Accelerating PoT Quantization on Edge Devices
Rappy Saha, Jude Haris, José Cano
Non-uniform quantization, such as power-of-two (PoT) quantization, matches data distributions better than uniform quantization, which reduces the quantization error of Deep Neural…