1 paper · 1 filter
Zimeng Wu, Donghao Wang, Chaozhe Jin +2
Long-context inference enhances the reasoning capability of Large Language Models (LLMs), but incurs significant computational overhead. Token-oriented methods, such as pruning and…