Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Ablation, Statistical Inference, and Validation for KV-Cache Compression
Paolo D'Alberto, Ashish Siarasao, Elliott Delaye +1
This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, through…
cs.LG2026
Statistical Inference and Quality Measures of KV Cache Quantisations Inspired by TurboQuant
Paolo D'Alberto
We analyse three KV cache quantization schemes under a fair bit budget: \textbf{KV} (scalar MSE baseline), \textbf{KQV} (WHT + MSE on ; WHT + MSE + QJL on ), and \textbf{QKQV…
cs.LG2024
Weight Block Sparsity: Training, Compilation, and AI Engine Accelerators
Paolo D'Alberto, Taehee Jeong, Akshai Jain +5
Nowadays, increasingly larger Deep Neural Networks (DNNs) are being developed, trained, and utilized. These networks require significant computational resources, putting a strain o…