collaborators

5 papers

cs.LG2026

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

Guangda Liu, Yiquan Wang, Chengwei Li +6

Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, mod…

cs.CV2026

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

Zhenyu Ning, Guangda Liu, Qihao Jin +4

Recent developments in Video Large Language Models (Video LLMs) have enabled models to process hour-long videos and exhibit exceptional performance. Nonetheless, the Key-Value (KV)…

cs.LG2026

FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference

Guangda Liu, Chengwei Li, Zhenyu Ning +5

Large language models (LLMs) are widely deployed with rapidly expanding context windows to support increasingly demanding applications. However, long contexts pose significant depl…

cs.LG2025

ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Guangda Liu, Chengwei Li, Jieru Zhao +2

Large Language Models (LLMs) have been widely deployed in a variety of applications, and the context length is rapidly increasing to handle tasks such as long-document QA and compl…

cs.GR2025

STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Support

Chenqi Zhang, Yu Feng, Jieru Zhao +4

3D Gaussian Splatting (3DGS) has gained popularity for its efficiency and sparse Gaussian-based representation. However, 3DGS struggles to meet the real-time requirement of 90 fram…