3 papers
cs.CL2024
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
Akide Liu, Jing Liu, Zizheng Pan +3
A critical approach for efficiently deploying computationally demanding large language models (LLMs) is Key-Value (KV) caching. The KV cache stores key-value states of previously g…
cs.LG2024
Efficient Stitchable Task Adaptation
Haoyu He, Zizheng Pan, Jing Liu +2
The paradigm of pre-training and fine-tuning has laid the foundation for deploying deep learning models. However, most fine-tuning methods are designed to meet a specific resource…
cs.CV2024
PaRa: Personalizing Text-to-Image Diffusion via Parameter Rank Reduction
Shangyu Chen, Zizheng Pan, Jianfei Cai +1
Personalizing a large-scale pretrained Text-to-Image (T2I) diffusion model is challenging as it typically struggles to make an appropriate trade-off between its training data distr…