3 papers
cs.DC2025
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
Qingyuan Li, Bo Zhang, Hui Kang +4
Nowadays, communication bottlenecks have emerged as a critical challenge in the distributed training and deployment of large language models (LLMs). This paper introduces FlashComm…
cs.AI2024
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
Qingyuan Li, Bo Zhang, Liang Ye +5
The ever-increasing sizes of large language models necessitate distributed solutions for fast inference that exploit multi-dimensional parallelism, where computational loads are sp…
eess.AS2020
AutoKWS: Keyword Spotting with Differentiable Architecture Search
Bo Zhang, Wenfeng Li, Qingyuan Li +3
Smart audio devices are gated by an always-on lightweight keyword spotting program to reduce power consumption. It is however challenging to design models that have both high accur…