1 paper
Robin Geens, Joran Heldens, Joren Dumoulin +1
Ternary weight quantization (e.g., BitNet b1.58) offers a promising path to mitigate the memory bandwidth bottleneck in Large Language Model (LLM) inference. However, conventional…