3 papers
cs.DC2026
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
Tiancheng Hu, Chenxi Wang, Ting Cao +9
Existing GPU-sharing techniques, including spatial and temporal sharing, aim to improve utilization but face challenges in simultaneously ensuring SLO adherence and maximizing effi…
cs.LG2025
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
Junhao Hu, Wenrui Huang, Weidong Wang +6
Large Language Models (LLMs) have demonstrated strong capabilities across various domains, with recent advancements in challenging reasoning tasks such as mathematics and programmi…
cs.LG2024
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
Junhao Hu, Wenrui Huang, Weidong Wang +7
Large Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become mor…