2 papers
cs.CL2026
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
Yixuan Wang, Qingyu Shi, Jiayu Zhou +3
The growing context length of Large Language Models (LLMs) enlarges the Key-Value (KV) cache, limiting deployment in resource-limited environments. Prior training-free approaches f…
cs.LG2024
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
Siyuan Huang, Yunchong Song, Jiayue Zhou +1
In the realm of graph learning, there is a category of methods that conceptualize graphs as hierarchical structures, utilizing node clustering to capture broader structural informa…