1 paper
Kaiyan Zhao, Tsuguchika Tabaru, Kenichi Kobayashi +3
Although recent quantized Large Language Models (LLMs), such as BitNet, have paved the way for significant reduction in memory usage during deployment with binary or ternary weight…