2 papers
cs.LG2026
HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference
Chun-Ting Chen, Dongmin Han, Hangyeol Mun +6
Block Quantization (BQ) is a promising approach for efficient deployment of large language models (LLMs), enabling low-precision computation with controlled accuracy degradation. C…
cs.AR2025
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
Chun-Ting Chen, HanGyeol Mun, Jian Meng +2
Edge inference for large language models (LLM) offers secure, low-latency, and cost-effective inference solutions. We emphasize that an edge accelerator should achieve high area ef…