1 paper · 1 filter
Wonpyo Park, Seung-won Hwang
Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods…