Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
Chuangtao Chen, Grace Li Zhang, Xunzhao Yin +3
Large Language Models (LLMs) rely heavily on Key-Value (KV) caching to minimize inference latency. However, standard KV caches are context-dependent: reusing a cached document in a…
cs.LG2026
Late Breaking Results: Conversion of Neural Networks into Logic Flows for Edge Computing
Daniel Stein, Shaoyi Huang, Rolf Drechsler +2
Neural networks have been successfully applied in various resource-constrained edge devices, where usually central processing units (CPUs) instead of graphics processing units exis…