1 citations · 1 across the 3 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
Haoxuan Wang, Chen Wang
The deployment of long-context Large Language Models (LLMs) poses significant challenges due to the intense computational cost of self-attention and the substantial memory overhead…
cs.DC2026
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
Xu Bai, Muhammed Tawfiqul Islam, Chen Wang +1
Pipeline parallelism (PP) is widely used to partition layers of large language models (LLMs) across GPUs, enabling scalable inference for large models. However, existing systems re…