2 papers
cs.LG2026
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
Yichun Xu, Navjot K. Khaira, Tejinder Singh
The key-value (KV) cache is a foundational optimization in Transformer-based large language models (LLMs), eliminating redundant recomputation of past token representations during…
cs.AI2025
AI Benchmark Democratization and Carpentry
Gregor von Laszewski, Wesley Brewer, Jeyan Thiyagalingam +28
Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring d…