3 papers
cs.DC2025
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
Qingyuan Li, Bo Zhang, Hui Kang +4
Nowadays, communication bottlenecks have emerged as a critical challenge in the distributed training and deployment of large language models (LLMs). This paper introduces FlashComm…
cs.AI2024
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
Qingyuan Li, Bo Zhang, Liang Ye +5
The ever-increasing sizes of large language models necessitate distributed solutions for fast inference that exploit multi-dimensional parallelism, where computational loads are sp…
cs.LG2024
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
Qingyuan Li, Ran Meng, Yiduo Li +5
We introduce Integer Scale, a novel post-training quantization scheme for large language models that effectively resolves the inference bottleneck in current fine-grained quantizat…