1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.DC2025
FailSafe: High-performance Resilient Serving
Ziyi Xu, Zhiqiang Xie, Swapnil Gandhi +1
Tensor parallelism (TP) enables large language models (LLMs) to scale inference efficiently across multiple GPUs, but its tight coupling makes systems fragile: a single GPU failure…
cs.DC2025★ 1 cited
Strata: Hierarchical Context Caching for Long Context Language Model Serving
Zhiqiang Xie, Ziyi Xu, Mark Zhao +5
Large Language Models (LLMs) with expanding context windows face significant performance hurdles. While caching key-value (KV) states is critical for avoiding redundant computation…