3 papers
cs.LG2026
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
Xindian Ma, Yidi Lu, Peng Zhang +1
The integration of visual information into Large Language Models (LLMs) has enabled Multimodal LLMs (MLLMs), but the quadratic memory and computational costs of Transformer archite…
cs.LG2025
The Scaling Law for LoRA Base on Mutual Information Upper Bound
Jing Zhang, Hui Gao, Peng Zhang +3
LoRA (Low-Rank Adaptation) is a widely used model fine-tuning method. In fine-tuning, the law among model performance, model parameters, and data complexity has been a focal issue…
cs.LG2024
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
Jing Zhang, Shuzhen Sun, Peng Zhang +5
Transformer-based large language models exhibit groundbreaking capabilities, but their storage and computational costs are prohibitively high, limiting their application in resourc…