2 papers
cs.LG2026
PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression
Zizhong Wang, Jieying Wang, Zhao Zhang +1
Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by r…
cs.DC2026
GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining
Jieying Wang, Shuyuan Fan, Mingkai Zheng +1
Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can signi…