2 papers
cs.CL2026
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
Jian Lin, Jiazhi Mi, Zicong Hong +5
Supporting long-context LLMs is challenging due to the substantial memory demands of the key-value (KV) cache. Existing offloading systems store the full cache in host memory and s…
cs.LG2024
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
Kuntai Du, Yuhan Liu, Yitian Hao +5
Deep learning inference on streaming media data, such as object detection in video or LiDAR feeds and text extraction from audio waves, is now ubiquitous. To achieve high inference…