1 paper
Chenyang Yin, Zhenyu Bai, Pranav Venkatram +3
Deploying Large Language Models (LLMs) efficiently on edge devices is often constrained by limited memory capacity and high power consumption. Low-bit quantization methods, particu…