3 papers
cs.DC2025
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
Qingyuan Li, Bo Zhang, Hui Kang +4
Nowadays, communication bottlenecks have emerged as a critical challenge in the distributed training and deployment of large language models (LLMs). This paper introduces FlashComm…
cs.CV2025
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
Bo Zhang, Shuo Li, Runhe Tian +4
In this paper, we introduce Flash-VL 2B, a novel approach to optimizing Vision-Language Models (VLMs) for real-time applications, targeting ultra-low latency and high throughput wi…
cs.AI2024
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
Qingyuan Li, Bo Zhang, Liang Ye +5
The ever-increasing sizes of large language models necessitate distributed solutions for fast inference that exploit multi-dimensional parallelism, where computational loads are sp…