64 citations · 97 across the 11 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
Weiye Wang, Chen Chen, Junxue Zhang +7
Distributed prefix caching has become a core technique for efficient LLM serving. However, for long-context requests with high cache hit ratios, retrieving reusable KVCache blocks…
cs.DC2019★ 23 cited
Quantifying the Performance of Federated Transfer Learning
Qinghe Jing, Weiyan Wang, Junxue Zhang +2
The scarcity of data and isolated data islands encourage different organizations to share data with each other to train machine learning models. However, there are increasing conce…