2 papers
cs.AR2025
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
Rohan Juneja, Shivam Aggarwal, Safeen Huda +2
Quantization is critical for efficiently deploying large language models (LLMs). Yet conventional methods remain hardware-agnostic, limited to bit-width constraints, and do not acc…
cs.AR2025
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
Chenyang Yin, Zhenyu Bai, Pranav Venkatram +3
Deploying Large Language Models (LLMs) efficiently on edge devices is often constrained by limited memory capacity and high power consumption. Low-bit quantization methods, particu…