Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ScaleBITS: Scalable Bitwidth Search for Hardware-Aligned Mixed-Precision LLMs
Xinlin Li, Timothy Chou, Josh Fromm +3
Post-training weight quantization is crucial for reducing the memory and inference cost of large language models (LLMs), yet pushing the average precision below 4 bits remains chal…
cs.LG2025
ICQuant: Index Coding enables Low-bit LLM Quantization
Xinlin Li, Osama Hanna, Christina Fragouli +1
The rapid deployment of Large Language Models (LLMs) highlights the need for efficient low-bit post-training quantization (PTQ), due to their high memory costs. A key challenge in…