2 papers
cs.LG2025
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
Junhao Hu, Wenrui Huang, Weidong Wang +6
Large Language Models (LLMs) have demonstrated strong capabilities across various domains, with recent advancements in challenging reasoning tasks such as mathematics and programmi…
cs.LG2025
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
Junhao Hu, Wenrui Huang, Weidong Wang +7
Large Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become mor…