activity
20242026
collaborators

7 papers

cs.CV2026

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping

Haotian Dong, Ye Li, Rongwei Lu +3

Visual autoregressive (AR) generation models have demonstrated strong potential for image generation, yet their next-token-prediction paradigm introduces considerable inference lat…

cs.LG2025

Taming Latency and Bandwidth: A Theoretical Framework and Adaptive Algorithm for Communication-Constrained Training

Rongwei Lu, Jingyan Jiang, Chunyang Li +2

Regional energy caps limit the growth of any single data center used for large-scale model training. This single-center training paradigm works when model size remains manageable,…

cs.CV2025

Accelerating Parallel Diffusion Model Serving with Residual Compression

Jiajun Luo, Yicheng Xiao, Jianru Xu +5

Diffusion models produce realistic images and videos but require substantial computational resources, necessitating multi-accelerator parallelism for real-time deployment. However,…

cs.DC2025

Staleness-Centric Optimizations for Parallel Diffusion MoE Inference

Jiajun Luo, Lizhuo Luo, Jianru Xu +4

Mixture-of-Experts-based (MoE-based) diffusion models demonstrate remarkable scalability in high-fidelity image generation, yet their reliance on expert parallelism introduces crit…

cs.DC2025

Beyond A Single AI Cluster: A Survey of Decentralized LLM Training

Haotian Dong, Jingyan Jiang, Rongwei Lu +5

The emergence of large language models (LLMs) has revolutionized AI development, yet the resource demands beyond a single cluster or even datacenter, limiting accessibility to well…

cs.LG2025

-FedHT: Stepsize-Aware Hard-Threshold Gradient Compression in Federated Learning

Rongwei Lu, Yutong Jiang, Jinrui Zhang +4

Gradient compression can effectively alleviate communication bottlenecks in Federated Learning (FL). Contemporary state-of-the-art sparse compressors, such as Top-, exhibit high…