papers
Publications (2)
cs.LG2026
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
Meng Xin, Sweta Priyadarshi, Jingyu Xin +26
This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-langu…
cs.LG2023
Tiered Pruning for Efficient Differentialble Inference-Aware Neural Architecture Search
SÅawomir Kierat, Mateusz Sieniawski, Denys Fridman +4
We propose three novel pruning techniques to improve the cost and results of inference-aware Differentiable Neural Architecture Search (DNAS). First, we introduce Prunode, a stocha…