#low-bit quantization

try —

5 papers match

cs.AR2026

LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

Sangjin Kim, Yuseon Choi, Jungjun Oh +2

LightRot introduces a lightweight rotation scheme and a dedicated hardware accelerator that enable energy‑efficient, low‑bit inference for large language models such as LLaMA2‑13B…

#low-bit quantization#large language models#hardware accelerator#fast hadamard transform
cs.AR2026

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference

Sangjin Kim, Yuseon Choi, Byeongcheol Kim +2

GyRot introduces a co-designed quantization framework and hardware accelerator that combine coarse rotation with fine-grained group quantization, enabling accurate 4-bit inference…

#low-bit quantization#large language models#rotation quantization#group quantization
cs.NE2026

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models

Ryona Noda

The paper proposes a joint optimization method for post‑training quantization of large language models that uses cross‑layer error compensation and finite‑sample feature‑statistics…

#post-training quantization#large language models#low-bit quantization#error compensation
cs.LG2026

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

Donghyun Son, Euntae Choi, Sungjoo Yoo

The paper proposes NSNQuant, a calibration‑free method that uses a double normalization and Hadamard transform to compress the key‑value cache of large language models with low‑bit…

#vector quantization#large language models#kv cache compression#low-bit quantization
cs.CV2026

Efficient Tuning Before Low-Bit Post-Training Quantization for Stochastic Gradient Descent-optimized Models

Peng Xia, Junbiao Pang, Muhammad Ayub Sabir

The paper introduces Efficient Tuning Before Quantization (ETBQ), a lightweight pre‑conditioning step that adjusts a full‑precision model using perturbations from quantization erro…

#post-training quantization#low-bit quantization#model tuning#sgd optimization

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.