3 papers
cs.LG2026
RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
Yuzong Chen, Xilai Dai, Jake Hyun +6
The recently introduced NVFP4 format demonstrates remarkable performance and memory benefits for quantized large language model (LLM) inference. However, we observe two types of re…
cs.LG2025
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
Yeonhong Park, Jake Hyun, Hojoon Kim +1
Quantization of Large Language Models (LLMs) has recently gained popularity, particularly for on-device settings with limited hardware resources. While efficient, quantization inev…
cs.DS2024
Log-Time K-Means Clustering for 1D Data: Novel Approaches with Proof and Implementation
Jake Hyun
Clustering is a key task in machine learning, with -means being widely used for its simplicity and effectiveness. While 1D clustering is common, existing methods often fail to e…