1 paper · 1 filter
Wei Shao, Lingchao Zheng, Pengyu Wang +3
Long context inference scenarios have become increasingly important for large language models, yet they introduce significant computational latency. While prior research has optimi…