13 citations · 13 across the 3 of their papers we have counts for
3 papers
cs.IR2026
MALLOC: Benchmarking the Memory-aware Long Sequence Compression for Large Sequential Recommendation
Qihang Yu, Kairui Fu, Zhaocheng Du +10
The scaling law, which indicates that model performance improves with increasing dataset and model capacity, has fueled a growing trend in expanding recommendation models in both i…
cs.LG2025★ 13 cited
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices
Zhaode Wang, Jingbang Yang, Xinyu Qian +4
Large language models (LLMs) have demonstrated exceptional performance across a variety of tasks. However, their substantial scale leads to significant computational resource consu…
cs.LG2025
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
Kunxi Li, Zhonghua Jiang, Zhouzhou Shen +5
This paper introduces MadaKV, a modality-adaptive key-value (KV) cache eviction strategy designed to enhance the efficiency of multimodal large language models (MLLMs) in long-cont…