Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Trainable Smooth-Rotation Transforms with Learned Channel Scales for LLM Quantization
Patrik Czakó, Gábor Kertész, Sándor Szénási
Post-training quantization (PTQ) is one of the most practical ways to reduce the serving cost of Large Language Models (LLMs), but activation quantization remains difficult because…
cs.LG2025
Turning LLM Activations Quantization-Friendly
Patrik Czakó, Gábor Kertész, Sándor Szénási
Quantization effectively reduces the serving costs of Large Language Models (LLMs) by speeding up data movement through compressed parameters and enabling faster operations via int…