2 papers
cs.DC2025
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
Tuowei Wang, Minxing Huang, Fengzu Li +3
As the demand for human-like reasoning, multi-turn dialogues, and long-form responses grows, large language models (LLMs) are increasingly expected to support efficient and effecti…
cs.MM2025
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
Deyu Zhang, Tingting Long, Jinrui Zhang +3
Enabling efficient text-video retrieval on edge-end devices is critical for real-world applications. Yet, existing methods face a critical challenge in balancing accuracy and compu…