Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
Chengzhu Bao, Xianglong Yan, Zhiteng Li +3
NVFP4 has recently emerged as an efficient 4-bit microscaling format for large language models (LLMs), offering superior numerical fidelity with native hardware support. However, e…
cs.LG2025
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
Yixin Song, Zhenliang Xue, Dongliang Wei +11
While frontier large language models (LLMs) continue to push capability boundaries, their deployment remains confined to GPU-powered cloud infrastructure. We challenge this paradig…