2 papers
cs.LG2026
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
Xiangyang Yin, Xingyu Liu, Tianhua Xia +5
Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increas…
cs.CL2025
SD: Self-Distilled Sparse Drafters
Mike Lasby, Nish Sinnadurai, Valavan Manohararajah +3
Speculative decoding is a powerful technique for reducing the latency of Large Language Models (LLMs), offering a fault-tolerant framework that enables the use of highly compressed…