1 paper · 1 filter
Suyu Ge, Xihui Lin, Yunan Zhang +2
Training and serving long-context large language models (LLMs) incurs substantial overhead. To address this, two critical steps are often required: a pretrained LLM typically under…