5 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CL2026
When Parallel Drafter Meets Parallel Speculative Decoding
Fuliang Liu, Xue Li, Kun Qian +4
DSpark-style parallel drafters have made speculative decoding highly effective, yet their draft phase remains serialized on the critical path of every round. Parallel speculative d…
cs.DC2025
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
Haoyu Chen, Xue Li, Kun Qian +3
In Large Language Model (LLM) inference services, it is challenging to make a parallelism strategy configuration, to efficiently process the requests of variance context lengths. R…
cs.DC2023★ 5 cited
Unicron: Economizing Self-Healing LLM Training at Scale
Tao He, Xue Li, Zhibin Wang +4
Training large-scale language models is increasingly critical in various domains, but it is hindered by frequent failures, leading to significant time and economic costs. Current f…