2 papers
cs.LG2026
RAP: Runtime Adaptive Pruning for LLM Inference
Huanrong Liu, Chunlin Tian, Xuyang Wei +2
Large language models (LLMs) excel at language understanding and generation, but their enormous computational and memory requirements hinder deployment. Compression offers a potent…
cs.IR2025
GREAT: Guiding Query Generation with a Trie for Recommending Related Search about Video at Kuaishou
Ninglu Shao, Jinshan Wang, Chenxu Wang +3
Currently, short video platforms have become the primary place for individuals to share experiences and obtain information. To better meet users' needs for acquiring information wh…