1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 1 cited
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
Weizhuo Li, Zhigang Wang, Yu Gu +1
Recently the generative Large Language Model (LLM) has achieved remarkable success in numerous applications. Notably its inference generates output tokens one-by-one, leading to ma…
cs.DC2024
LR-CNN: Lightweight Row-centric Convolutional Neural Network Training for Memory Reduction
Zhigang Wang, Hangyu Yang, Ning Wang +5
In the last decade, Convolutional Neural Network with a multi-layer architecture has advanced rapidly. However, training its complex network is very space-consuming, since a lot of…
cs.DC2024
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
Zhigang Wang, Xu Zhang, Ning Wang +5
Transformer-based models are becoming deeper and larger recently. For better scalability, an underlying training solution in industry is to split billions of parameters (tensors) i…