3 papers
cs.CV2026
SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers
Wonsuk Jang, Thierry Tambe
Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder edge deployment. Quantization can…
cs.LG2026
RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
Yuzong Chen, Xilai Dai, Jake Hyun +6
The recently introduced NVFP4 format demonstrates remarkable performance and memory benefits for quantized large language model (LLM) inference. However, we observe two types of re…
cs.CL2025
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
Wonsuk Jang, Thierry Tambe
The rapidly increasing size of large language models (LLMs) presents significant challenges in memory usage and computational costs. Quantizing both weights and activations can add…