6 papers
KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
Chen Qiu, Ziwu Liu, Chao Fei +2
KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We present KVDiagnosis, a diagnostic dataset and b…
ART: Attention Run-time Termination for Efficient Large Language Model Decoding
Chen Qiu, Guozhong Li, Cristian McGee +2
Long-context decoding in Large Language Models (LLMs) is constrained by the cost of accessing and processing the Key-Value (KV) cache. Despite evidence that attention outputs depen…
Can Deep Neural Networks Improve Compression of Very Large Scientific Data?
Muhannad Alhumaidi, Guozhong Li, Spiros Skiadopoulos +1
Error-bounded lossy compression is a fundamental technique for managing the rapidly growing volumes of scientific data produced by modern simulations and observational instruments.…
CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference
Chao Fei, Guozhong Li, Chenxi Liu +1
Long-context LLMs demand accurate inference at low latency, yet decoding becomes primarily constrained by KV cache as context grows. Prior pruning methods are largely context-agnos…
LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
Guozhong Li, Muhannad Alhumaidi, Spiros Skiadopoulos +1
The rapid growth of high-resolution scientific simulations and observation systems is generating massive spatiotemporal datasets, making efficient, error-bounded compression increa…
GraphComp: Extreme Error-bounded Compression of Scientific Data via Temporal Graph Autoencoders
Guozhong Li, Muhannad Alhumaidi, Spiros Skiadopoulos +2
The generation of voluminous scientific data poses significant challenges for efficient storage, transfer, and analysis. Recently, error-bounded lossy compression methods emerged d…