2 papers
cs.LG2026
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
Kabir Swain, Sijie Han, Daniel Karl I. Weidele +3
We propose \textbf{Hurwitz Quaternion Multiplicative Quantization (HQMQ)}, a \textbf{calibration-free} method for KV cache compression of large language models. HQMQ treats each 4-…
cs.DC2025
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
Michael V. DeBole, Rathinakumar Appuswamy, Neil McGlohon +30
A vertically integrated, end-to-end, research prototype system combines 288 NorthPole neural inference accelerator cards, offline training algorithms, a high-performance runtime st…