2 papers
cs.DC2026
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
Wanqian Li, Jintao Peng, Zongfei Jing +7
Large language model (LLM) inference increasingly depends on multi-GPU execution, yet existing inference parallelization strategies require layer-wise inter-rank synchronization, m…
cs.AR2024
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
Xiaotian Sun, Xinyu Wang, Wanqian Li +2
Various processing-in-memory (PIM) accelerators based on various devices, micro-architectures, and interfaces have been proposed to accelerate deep neural networks (DNNs). How to d…