2 papers
cs.LG2026
Multi-Head Low-Rank Attention
Songtao Liu, Hongwu Peng, Zhiwei Zhang +2
Long-context inference in large language models is bottlenecked by Key--Value (KV) cache loading during the decoding stage, where the sequential nature of generation requires repea…
cs.CL2024
MAPO: Boosting Large Language Model Performance with Model-Adaptive Prompt Optimization
Yuyan Chen, Zhihao Wen, Ge Fan +6
Prompt engineering, as an efficient and effective way to leverage Large Language Models (LLM), has drawn a lot of attention from the research community. The existing research prima…