9 papers
Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models
Jinyang Du, Shenghao Jin, Ziqian Xu +5
Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a large resident parameter footprint…
DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving
Hantian Zha, Teng Ma, Yang Yong +7
Diffusion-based generation is increasingly powering production content pipelines; however, deploying these models at scale remains a significant challenge. Model weights frequently…
QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks
Haotong Qin, Xudong Ma, Xianglong Liu +4
Low-bit quantization is widely used to compress super-resolution (SR) models and reduce storage and computation costs for deployment on resource-limited devices. However, when SR m…
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
Yushi Huang, Zining Wang, Zhihang Yuan +5
Mixture-of-Experts (MoE) Multimodal large language models (MLLMs) excel at vision-language tasks, but they suffer from high computational inefficiency. To reduce inference overhead…
Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge
Li Kang, Heng Zhou, Xiufeng Song +41
Recent advancements in multimodal large language models and vision-languageaction models have significantly driven progress in Embodied AI. As the field transitions toward more com…
Past-Future Scheduler for LLM Serving under SLA Guarantees
Ruihao Gong, Shihao Bai, Siyu Wu +5
The exploration and application of Large Language Models (LLMs) is thriving. To reduce deployment costs, continuous batching has become an essential feature in current service fram…