Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding
Chao Fang, Jun Yin, Man Shi +1
With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this ch…
cs.AR2024
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
Chao Fang, Man Shi, Robin Geens +3
The widely-used, weight-only quantized large language models (LLMs), which leverage low-bit integer (INT) weights and retain floating-point (FP) activations, reduce storage require…