2 citations · 4 across the 10 of their papers we have counts for
4 papers · 1 filter
Ablation, Statistical Inference, and Validation for KV-Cache Compression
Paolo D'Alberto, Ashish Siarasao, Elliott Delaye +1
This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, through…
Statistical Inference and Quality Measures of KV Cache Quantisations Inspired by TurboQuant
Paolo D'Alberto
We analyse three KV cache quantization schemes under a fair bit budget: \textbf{KV} (scalar MSE baseline), \textbf{KQV} (WHT + MSE on ; WHT + MSE + QJL on ), and \textbf{QKQV…
Weight Block Sparsity: Training, Compilation, and AI Engine Accelerators
Paolo D'Alberto, Taehee Jeong, Akshai Jain +5
Nowadays, increasingly larger Deep Neural Networks (DNNs) are being developed, trained, and utilized. These networks require significant computational resources, putting a strain o…
Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
Sean O. Settle, Manasa Bollavaram, Paolo D'Alberto +6
Deep learning as a means to inferencing has proliferated thanks to its versatility and ability to approach or exceed human-level accuracy. These computational models have seemingly…