2 papers
cs.LG2026
FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy
Qingyao Yang, Runming Yang, He Xiao +7
While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specializ…
cs.LG2025
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
He Xiao, Qingyao Yang, Dirui Xie +7
Large language models with billions of parameters are often over-provisioned: many layers contribute little unique information yet dominate the memory and energy footprint during i…