3 papers
cs.NI2026
R2E-VID: Two-Stage Robust Routing via Temporal Gating for Elastic Edge-Cloud Video Inference
Zheming Yang, Lulu Zuo, Shun Lu +4
With the rapid growth of large-scale video analytics applications, edge-cloud collaborative systems have become the dominant paradigm for real-time inference. However, existing app…
cs.DC2025
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
Ziming Liu, Boyu Tian, Guoteng Wang +15
Mixture-of-Experts (MoE) models challenge serving infrastructures with dynamic, sparse expert utilization, causing instability on conventional systems designed for dense architectu…
cs.GR2025
SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation
Shenggan Cheng, Yuanxin Wei, Lansong Diao +8
Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing t…