2 papers
cs.IR2025
Time Matters: Enhancing Sequential Recommendations with Time-Guided Graph Neural ODEs
Haoyan Fu, Zhida Qin, Shixiao Yang +5
Sequential recommendation (SR) is widely deployed in e-commerce platforms, streaming services, etc., revealing significant potential to enhance user experience. However, existing m…
cs.LG2024
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
Yanqi Zhang, Yuwei Hu, Runyuan Zhao +2
Large language models (LLMs) demonstrate remarkable capabilities but face substantial serving costs due to their high memory demands, with the key-value (KV) cache being a primary…