2 citations · 2 across the 5 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving
Zikun Li, Yixuan Mei, Shiqi Pan +9
LLM serving systems increasingly disaggregate inference into finer-grained stages, with recent approaches separating attention from FFN or MoE execution during decode. This operato…
cs.DC2026
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
Yixuan Mei, Zikun Li, Zixuan Chen +5
The usage of large language models (LLMs) has grown increasingly fragmented, with no single model dominating. Meanwhile, cloud providers offer a wide range of mid-tier and older-ge…