2 papers
cs.DC2025
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Wentao Liu, Yuhao Hu, Ruiting Zhou +2
Mixture-of-Experts (MoE) has become a dominant architecture in large language models (LLMs) due to its ability to scale model capacity via sparse expert activation. Meanwhile, serv…
cs.CV2025
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
Youping Gu, Xiaolong Li, Yuhao Hu +2
Diffusion Transformers currently lead the field in high-quality video generation, but their slow iterative denoising process and prohibitive quadratic attention costs for long sequ…