1 paper
Zi-Wei Lin, Tian-Sheuan Chang
Deploying Large Language Models (LLMs) on resource-constrained edge devices faces critical bottlenecks in memory bandwidth and power consumption. While ternary quantization (e.g.,…